Skip to content

My anycast nodes diverged for 4 hours—nobody noticed

Networking by Nadia 2 replies 175 views
#1

| 2025-04-11 06:14 UTC: OVHcloud PoP in RackNerd region starts serving stale zone
| 2025-04-11 06:14 UTC: Hetzner PoP continues correct zone
| 2025-04-11 10:22 UTC: I notice during unrelated dig query
| 2025-04-11 10:25 UTC: EOF (fixed)

The monitoring? Redirected to /dev/null apparently. Health checks were piping to the wrong endpoint since a config push Tuesday. Users in APAC got one answer, EMEA got another. Some TTLs meant cached wrongness for hours after.

I have since piped alerts to three independent channels. No more single points of redirection.

EOF

/dev/null: full of good ideas
#2

4 hrs? Back in my day we had split-brain for 3 days on a VAX cluster and the users just thought the net was "slow" :-)

U r lucky anyone noticed at all with modern attention spans

2 #3

Oh honey, four hours of divergent DNS is a backup nightmare waiting to happen. Did you test your restore after you fixed it?

And what's your 3-2-1 rule looking like for those zone files?

  • 3 copies minimum?
  • 2 different media types?
  • 1 offsite?

I hope so. I've seen people "fix" split-brain and then discover their secondary was the only one with clean data.

3-2-1 or you're already dead

Post a reply

You need an account to reply. Log in or register to join the conversation.

Post reply Preview Save draft