Skip to content

Authoritative DNS—when is self-hosting masochism?

VPS Hosting by blogfranck 3 replies 232 views
#1

I have been running my own PowerDNS cluster on three HostHatch VPS for two years now. The hosting costs are minimal, the control is absolute. Yet I read everywhere that Cloudflare is "the only sensible choice" for Authoritative DNS. At what point does self-hosting become mere masochism?

My setup: three nodes, anycast-lite with BGP on KnownHost, daily AXFR to Hetzner cold standby. I monitor with Prometheus. It works. It is elegant. Voilà.

But I wonder—am I one DDoS away from catastrophe? The Dyn attack of 2016 taught us something, no? I was not in hosting then. Those who were, please share your wisdom. I do not wish to be the fool who learns only by fire.

#2

Ngl running your own auth dns for fun is one thing but for actual production traffic? Lol

I had a similar setup for my llm inference cluster coordinator, like 8x A100s talking to each other, every dns hiccup = vram context loss = $$$ burning. Switched to anycast, never looked back

Your "minimal cost" ignores the incident response cost when you get woken up at 3am because someone's reflection attacking udp/53

CUDA cores are my love language
#3

Actually I was there during the Dyn attack and I can tell you that the Problem was not actually self-hosting per se but actually the monoculture of using only one Provider or? We had our own BIND infrastructure actually running on Leaseweb with proper rate-limiting and query-patternsanalysis and actually we survived because the AttackSurface was distributed across multiple networklocations with different upstreamproviders or? The ones who actually suffered were those who actually relied on a single SaaS without actually understanding how the DNSresolution actually worked under the hood or? So actually your PowerDNS cluster is actually not automatically foolish but actually you need actually proper monitoring and actually responseplaybooks and actually maybe some automatedblackholing or?

But actually Cloudflare has actually better anycast and actually more bandwidth than your three HostHatch instances actually combined or?

#4

Verify your restores

I was at a shop that ran their own auth dns on "premium" dedis. 2016 came, 1.2 Tbps reflection flood, upstream null-routed the /24 before we even got paged. Took 6 hours to get anycast reannounced. 6 hours of email not flowing. Verify your restores weekly, people.

Your cold standby on Hetzner is worthless if the attack follows your delegation. Have you TESTED failing over during simulated packet loss? Have you tested your restore? I used https://www.hetzner.com/cloud for cold standby.

Paranoid? I was called paranoid in 2015. In 2016 I was called employed.

Post a reply

You need an account to reply. Log in or register to join the conversation.

Post reply Preview Save draft