Skip to content

When the datacenter lost power — your stories from the dark

Datacenter Talk by uma 27 replies 2.9K views
#1

I will start. December 2024. Contabo facility outside St. Louis. Their status page went green at 09:00. My external monitoring showed 100% packet loss from 08:47 to 09:00. Thirteen minutes of darkness that never existed on their status page. What happened: utility feed failed, transfer switch to generator did not engage. UPS carried the load for 8 minutes, then batteries hit critical and the whole floor went dark. My VPS on a Ryzen 5950X node, 18 months uptime, gone in a blink. They had a generator. It worked. The transfer switch did not. Classic single point of failure. I learned about it from my phone alerts before their status page updated. Alert fatigue is real but this is why I run three external monitors. 99.99% or bust? I got 99.997% for the year. The gap between that and perfect is where the stories live.

436 days. reboot is surrender.
#2
uma said:
Transfer switch to generator did not engage

Contabo St. Louis, I know that building. Used to ship pallets there. 42U racks, dual 30A feeds, but the transfer switches were Eaton units from 2016. Those contactors wear out. I have replaced six of them.

My story: OVHcloud facility in Hillsboro, July 2024. 110F outside. Utility brownout, generators start fine, but UPS battery string 3B had a thermal runaway domino. One battery cracked, vented, took out the string next to it. By the time the BMS isolated, we had lost 40% of UPS capacity. The remaining strings hit discharge limit in 4 minutes instead of 10.

I had a Supermicro 2029P there, 2x Xeon Gold 6248R, 512GB RAM, 1.2kW at the rail. That machine cost me $4,200 used plus $85/month for the half rack. For that money you get old enterprise gear and you pray the DC did their battery impedance testing. They had not. I found the maintenance logs later. Last test was 14 months prior.

The smell was the worst part. Sulfur and hot plastic. You don't forget it.

visit twice: install and decom
#3
  • Facility: Hetzner, Ashburn region, February 2025.
  • Incident duration: 31 hours, 17 minutes.
  • Root cause: diesel fuel gelling at -22C, fuel truck stuck on highway 401 for 14 hours.
  • Affected systems: 4x KVM VPS nodes (E5-2680v4, 256GB RAM each), 1x dedicated Ryzen 9 7950X box.

The facility had three 2MW generators with 48-hour belly tanks. They ran for 34 hours before the fuel truck arrived. During hour 19, generator 2 threw a coolant alarm and auto-shutdown. The remaining two generators ran at 85% load, which is above the 80% recommended continuous rating per Caterpillar C32 datasheet.

My monitoring stack:

  • Prometheus 2.48.0 with node_exporter 1.7.0
  • External: UptimeRobot + self-hosted Uptime Kuma 1.23.0
  • Alertmanager with PagerDuty integration

I received 47 alerts. I acknowledged 12. The rest were duplicates or flapping during the generator switch. Alert fatigue is measurable. I have the data.

It's always DNS. Always.
#4

We had a Nagios box that emailed a Blackberry. When the power went out, you drove to the datacenter with a flashlight and a sweater because the HVAC was on the same feed and it got cold fast.

Kids these days with their 47 alerts and their mean time to acknowledge. In 2009 I watched a facility in New Jersey run on portable generators for two days. Not building generators. Portables. 500kW tow-behind units daisy-chained with camlock cables running through a loading dock door propped open with a 2x4. The fuel truck came every six hours. The loading dock froze open. We lost two drives from condensation when they finally brought the building back online.

Mark my words, the next big outage will be someone trusting their cloud availability zone and learning that "redundant" does not mean "independent."

IPv4, IRC, and irssi — fight me
#5

Status page green at 09:00 but your graphs say 08:47. Classic.

#6

Which fuel additive did Hetzner use, if any? -22C is rough.

#7

Two days on portables.

#8
ClearIce said:
Which fuel additive did Hetzner use, if any?

None that I know of. The fuel spec was #2 diesel with cloud point around -12C. At -22C ambient, straight #2 gels without additive or blended #1 kerosene. I do not know if the facility had pre-treated fuel or if they assumed the belly tanks would stay warm enough from generator waste heat. Generator 2's coolant alarm was unrelated, cracked hose, but it cost us redundancy.

I have the Prometheus graphs. Load on remaining gens climbed from 62% to 85% in about 90 seconds after gen2 shut down. You could watch it step up on each automatic load shed failure. They never dropped non-critical.

It's always DNS. Always.
#9
FlowSana said:
They never dropped non-critical.

Because nobody wants to be the guy who explains why the billing database went down to keep the badge readers warm. Load shedding plans look great on paper. First time you actually execute one, some VP is on a plane and suddenly his VPN concentrator is "critical infrastructure." I have seen it.

The New Jersey portables were at a facility that is now a WeWork. The camlock cables froze to the loading dock because snow blew in. We had to pour hot water on them to break them free. Hot water. On electrical connections. OSHA would have loved us.

IPv4, IRC, and irssi — fight me
#10
haroldgsm said:
Nobody wants to be the guy who explains why the billing database went down

This is why I run three status page monitors plus my own. Not because I do not trust the provider. Because I do not trust the provider's incentives to be honest about the severity. Thirteen minutes of 100% packet loss and Contabo's page says "degraded connectivity." Degraded. My monitoring says DEAD. Theirs says "some customers may experience latency." The gap between those two statements is why I have alert fatigue.

I moved my main workload to Hetzner Nuremberg after St. Louis. Nuremberg has had two incidents in three years, both under ten minutes, both acknowledged within four minutes. That is a track record I can sleep with.

436 days. reboot is surrender.

Post a reply

You need an account to reply. Log in or register to join the conversation.

Post reply Preview Save draft