Skip to content

When the datacenter lost power — your stories from the dark

Datacenter Talk by uma 27 replies 2.9K views
#11

Uma you ever think the status page is green because the person who updates it is on the same power feed as the servers

8 #12
uma said:
Moved my main workload to Hetzner Nuremberg

Nuremberg is solid. I have hardware there, SX134 dedicated, been running since 2021. One planned maintenance, 23 minutes, announced two weeks ahead. The Ashburn and Hillsboro facilities are newer and it shows in the rough edges.

The Eaton transfer switches at Contabo St. Louis, those contactors have a rated mechanical life of 100,000 operations but electrical life is maybe 15,000 under load. If they test monthly, that is 12,000 operations in ten years. You are dancing on the edge. I bet they never counted test cycles against wear.

visit twice: install and decom
#13

I am in Toulouse and I use OVHcloud Gravelines. Never lost power but I lost network twice in 2023. Their status page was accurate both times, to be fair. Ten minutes, fifteen minutes, "network degradation" and it was actually just degradation, not dead.

The thing with diesel gelling, is that not a solved problem in Canada? I thought all northern facilities ran winter blend by default.

#14
moetrain said:
Is that not a solved problem in Canada?

You would think. The facility is in Ashburn, Virginia, not Toronto. The -22C was an arctic blast that hit the US northeast in February 2025. Toronto was -15C that week, which is normal for us. The fuel truck was stuck on I-401 in Ontario trying to cross to the US, not in Virginia. The fuel came from a Canadian supplier because the US terminal was also backed up from the same storm.

I have the ticket emails. "Fuel delivery delayed due to weather." Hour 12. Hour 18. Hour 24. Then "generator 2 coolant alarm." Then silence for 40 minutes while they found the hose and rerouted.

It's always DNS. Always.
#15

I am in Jakarta and I use Contabo Singapore. My uptime is 99.7% over two years which sounds good until you do the math. 99.7% is 26 hours down per year. I get packet loss every rainy season. Their status page never shows it. I think the undersea cable to Singapore has issues but Contabo blames "network maintenance" every time.

I pay 8 EUR for 4 cores and 8GB — https://contabo.com/en/vps/. At that price I do not complain much. But I also do not host anything that pays my rent.

#16
RanggaDass said:
I pay 8 EUR for 4 cores and 8GB

Same plan, same problems, but I am in Bandung so the latency to Singapore is better than yours probably. I get 18ms, you probably get 35ms? I also see the rainy season packet loss. It is not Contabo, it is the Telin cable. You can check with BGP looking glass from SGIX. The loss is on the path, not the host.

I have three Contabo VPS and one Hetzner cloud in Helsinki. The Hetzner one has never had an unplanned minute in 14 months. But it costs 14 EUR for less RAM. You get what you pay for.

#17

I am in Stockholm and I use Hetzner Falkenstein. 12ms to Frankfurt, 180ms to everywhere I actually need, which is why I also have a VM in Helsinki now. The Falkenstein facility had a cooling incident in 2022, not power, but they posted full details within two hours. I respect that.

The diesel gelling story surprises me. Hetzner is usually better at winter ops. Their Finnish datacenter should have taught them.

1ms or I don't want it
#18

I am in Tallinn and I use Hetzner Helsinki. 5ms away, beautiful. One planned maintenance in 18 months, 8 minutes, announced 10 days ahead. I have a Ryzen 9 7950X dedicated there for game server hosting. The players notice 8 minutes.

I also have a Contabo VPS in Nuremberg for backups. It has been fine but I do not trust it for production after reading this thread.

#19
liao_scope said:
I do not trust it for production after reading this thread

Smart. Trust is not about the brand. It is about the specific building, the specific maintenance contract on the specific generator, the specific night crew who may or may not know how to manually throw a transfer switch at 3AM while the UPS is screaming. I have met those night crews. Some are heroes. Some are two weeks out of a Best Buy Geek Squad and the only reason they are there is because the real engineer called in sick.

The next big outage will be a facility that passed every audit, every generator test, every thermal scan. And then the one guy who knew how to read the BMS alarm codes retired and nobody updated the runbook.

IPv4, IRC, and irssi — fight me
#20

This thread is why I keep a cold spare at a different provider in a different city. My main is Hetzner Falkenstein, spare is OVHcloud London. The spare costs me 45 EUR a month to do nothing. I have failed over to it twice in three years, both times for my own mistakes, not provider outages. But I sleep better.

Has anyone actually tested their failover for real? Not a drill. Actual "production is down, button now." The first time I did it, my DNS TTL was 300 seconds and my CDN cached the error page for 24 hours. Took me two hours to realize half my users still saw the broken site.

Post a reply

You need an account to reply. Log in or register to join the conversation.

Post reply Preview Save draft