Skip to content

When the datacenter lost power — your stories from the dark

Datacenter Talk by uma 27 replies 2.9K views
#21
ServersJohn said:
My DNS TTL was 300 seconds and my CDN cached the error page for 24 hours

This. This right here. Your infrastructure is only as good as your slowest cache. I learned this with Cloudflare page rules. You can have a beautiful health-checked anycast setup and one stale edge cache in Singapore serves your 500 error to all of APAC for an hour.

I use Hetzner and OVHcloud both. Not for redundancy, just for price shopping. Hetzner is cheaper for CPU, OVHcloud is cheaper for bandwidth. Neither has ever died on me but I have never hosted anything that pays my rent either.

#22

Following this thread. I have a single Contabo VPS in St. Louis and now I am nervous.

#23
green15 said:
Now I am nervous

Good. Nervous is correct. But also practical: what is your actual downtime cost? My St. Louis VPS was 4 EUR a month for a testbed. The downtime cost me nothing but annoyance. I kept it running because I am lazy, not because it mattered. The stuff that matters is in Nuremberg and I pay 45 EUR there for better redundancy.

If your VPS is a personal blog, keep it, sleep fine. If it is your business, have a plan. The plan does not have to be expensive. It has to be tested.

436 days. reboot is surrender.
3 #24

Has anyone used the Hetzner storage box for backups during an outage? I am wondering if the storage network is on the same power feeds as the compute. My storage box in Falkenstein stayed up during the 2022 cooling incident but I do not know if that was luck or design.

#25
tier_ethqn said:
Same power feeds as the compute

Different substations in Falkenstein, or so they told me in a ticket. The storage boxes are in a separate building. But "separate building" can still mean "same fence, same fuel truck" if the logistics contract is centralized. I do not know their fuel supplier arrangement.

I back up to Wasabi from Hetzner. Different company, different continent, different failure domain. Costs me 7 USD a month for 500GB. Paranoid? Maybe. But I have restored from it twice.

It's always DNS. Always.
#26

I run Nagios still. On a Pi in my closet. It emails me. When my closet loses power, Nagios stops emailing. I know this is stupid. I am posting in this thread because I know this is stupid and I still have not fixed it.

#27
systems52 said:
I know this is stupid

Self-awareness is the first step. The second step is a UPS for the Pi, which costs 80 dollars and gives you 45 minutes. The third step is a second Pi at your mother's house, which costs her nothing and gives you off-site monitoring. The fourth step is admitting you will never do steps two or three and just paying for UptimeRobot like a normal person.

I have been at step four for six years. UptimeRobot emails my phone. My phone has battery. Problem solved for 8 dollars a month.

IPv4, IRC, and irssi — fight me
#28

I am in Manchester and I use a dedi at Hetzner Falkenstein plus a VPS at OVHcloud London. My marriage lasted 8 years, my server uptime is at 6 years and counting. The server is more stable. I have failed over once, OVHcloud London to Falkenstein, when a roadworks crew cut fiber outside the London datacenter. OVHcloud status page said nothing for 22 minutes. My monitoring caught it at minute 1. Their Twitter account posted before their status page updated.

Status pages are a marketing tool. Treat them as such.

Post a reply

You need an account to reply. Log in or register to join the conversation.

Post reply Preview Save draft