Two DNS Layers That Look Identical When They're Both Green

aws route53 dns vpc networking troubleshooting

Pods in one account’s VPC kept getting NXDOMAIN for names in a private hosted zone owned by another account. The failure had a very specific shape: it only happened from that one foreign VPC. Query the same names from the zone’s home VPC and they resolved instantly. So this was not a broken zone, not a bad record, not a propagation delay. It was scoped to exactly one VPC, which is the kind of detail that should have told me what was wrong and instead told me nothing, because I was reading it through the wrong layer.

The peering connection between the two VPCs was healthy. Routes were in place on both sides. Security groups allowed the traffic. And the peering connection’s allow_remote_vpc_dns_resolution option was checked, which is the setting whose name most sounds like “let this VPC resolve DNS across the peering.” Every signal I knew to look at was green. The names still would not resolve.

The two layers

The reason none of those green signals helped is that DNS across a peered, cross-account setup runs on two independent mechanisms, and they look the same from the outside.

The first is the peering connection’s DNS resolution. That allow_remote_vpc_dns_resolution flag does one narrow thing: it lets instances on one side of the peering resolve the provider-assigned private DNS names of instances on the other side, the ip-10-x-x-x.region.compute.internal hostnames AWS hands out automatically. That is the entire scope of the flag. It is about the default internal DNS that every VPC gets for free.

The second is private hosted zone membership. A private hosted zone answers queries only from the VPCs that are associated with it. Association is a distinct, explicit act: you attach a specific VPC to a specific zone, and only then does that zone appear in that VPC’s resolver. A zone that is not associated with your VPC does not exist as far as your VPC’s resolver is concerned, and the honest, correct answer to a query for one of its names is NXDOMAIN.

These two layers have nothing to do with each other. Peering DNS resolution is about the automatic compute.internal names; zone membership is about your custom zone. Checking the peering box changes the first and does absolutely nothing to the second. My foreign VPC could resolve the other side’s compute.internal hostnames perfectly, which is exactly what the flag promises, and that success told me nothing about whether my custom zone would answer, because that was never the flag’s job.

Why it fools you

Every diagnostic reflex I had pointed at the peering layer. Is the peering active? Are the routes there? Do the security groups allow it? Is DNS resolution enabled on the connection? Those are the questions you ask about cross-VPC connectivity, and every single answer came back correct, because the peering layer genuinely was correct. The problem was that a correct peering layer is not evidence about zone membership. I kept collecting green checks from one mechanism and treating them as reassurance about a different one.

Nothing in that whole surface says “this VPC is not a member of the zone.” There is no red light for it. The peering console is happy, the route tables are happy, the security groups are happy, and the missing association lives in a completely different resource that you only find if you already suspect it. The failure sits in the gap between two layers that both report green, and the shape of the bug is that the green from the layer you can see convinces you the layer you can’t see is fine too.

The fix, and why it is two steps

The fix is to make the foreign VPC an actual member of the zone. Because the zone and the VPC live in different accounts, that association is a two-step handshake, one call in each account:

# In the zone-owner account: authorize the foreign VPC.
aws route53 create-vpc-association-authorization \
  --hosted-zone-id "$ZONE_ID" \
  --vpc VPCRegion="$REGION",VPCId="$FOREIGN_VPC_ID" \
  --profile zone-owner

# In the VPC-owner account: accept by associating.
aws route53 associate-vpc-with-hosted-zone \
  --hosted-zone-id "$ZONE_ID" \
  --vpc VPCRegion="$REGION",VPCId="$FOREIGN_VPC_ID" \
  --profile vpc-owner

The split is not ceremony. The authorization is the zone owner’s decision to make, so it runs under the zone owner’s credentials; the association attaches the zone to a VPC that belongs to the other account, so it runs under that account’s credentials. Each half runs where its resource lives. Run them in that order, and the foreign VPC’s resolver starts answering for the zone immediately. The authorization record persists after the association completes, so both halves are durable and, later, both are importable into whatever manages your infrastructure as code.

The audit trap

Once you know association is the thing that matters, there is a second surprise waiting: there is no single command that tells you whether a VPC can resolve a zone. There are two, and they answer different questions.

# What is actually associated (the resolver truth):
aws route53 get-hosted-zone --id "$ZONE_ID"

# What has merely been authorized (a standing invitation):
aws route53 list-vpc-association-authorizations --hosted-zone-id "$ZONE_ID"

get-hosted-zone lists the VPCs actually associated, which is the set that can resolve. list-vpc-association-authorizations lists the VPCs that have been authorized, which is only permission to associate, not the association itself. They routinely disagree. An audit of one zone turned up a VPC that had been authorized but never associated: from the infrastructure-as-code side it looked done, the authorization resource existed, and resolution from that VPC was still broken. The reverse also happens: the same zone carried an association to a VPC that no longer existed in any account, a leftover from a torn-down environment, harmless but easy to mistake for a live peer. Authorized is not associated, and associated is not necessarily still real.

The lesson worth keeping

The durable lesson is not about Route53. It is that when two layers present the same green light, a healthy signal from one is not evidence about the other, no matter how confidently its name implies otherwise. allow_remote_vpc_dns_resolution sounds like it governs all DNS across the peering. It governs one narrow slice, and the slice I actually needed was in a different resource entirely.

So when a signal is green and the thing still does not work, the useful question is not “why is this green thing lying to me.” The green thing is usually telling the truth about its own layer. The question is which layer the failure actually lives in, and whether the reassuring green you are staring at has any authority there at all. Most of the time it does not, and the real answer is one resource over, in a place you never looked because nothing pointed you there.

$ cat KUBERNETES .md
· 8 min read

The EKS 401 That Wasn't About Credentials

Our CI pipeline was minting a valid EKS token, presenting it cleanly, and getting rejected. The problem was not AWS creds. It was EKS access control, and the failure was hiding in plain sight in the cluster's RBAC.

kubernetes eks aws devops platform-engineering troubleshooting