Skip to content

Your worst outage story — go

General Discussion by uma 13 replies 875 views
#1

I'll start. November 2nd, 2024. My primary monitoring node at Contabo went dark at 14:03 UTC. 36 hours, 14 minutes. Not a blip, not a taper. Hard vertical drop to zero. I had three layers: pingdom-style externals, a status page on OVHcloud, and a self-hosted Uptime Kuma instance. All three screaming within 90 seconds. Alert fatigue is real — I had 47 notifications by hour six. The provider's own status page? Green the entire time. "Scheduled maintenance" posted at hour 22, backdated. I have screenshots. The root cause was a core router firmware bug. They didn't have a rollback image. I learned that my 99.99% SLA was calculated per-calendar-year, so one 36-hour hit in November still let them claim "annual target under review." I migrated everything to Hetzner https://www.hetzner.com with dual-stack BGP by December. 99.997% since

436 days. reboot is surrender.
#2
uma said:
36 hours, 14 minutes

36 hours on a single router with no rollback. That's amateur hour.

My worst was 2019. Shipped a pallet of Dell R630s to a colo in Kansas City. 8 servers, 2U each, dual E5-2680v3, 256GB RAM per box. Cross-country freight, $3400 just to move the metal.

They go live on a Thursday. Saturday 2AM the facility loses UPS B. Transfer switch fails. My whole rack goes dark. The kicker? I had RAID 10 on every box, battery-backed, the works. Didn't matter. When power came back, 6 of 8 RAID cards bricked mid-rebuild. LSI 9361-8I, known firmware flaw after unclean shutdown.

48 hours to source replacement cards. Another 36 to rebuild arrays. Total downtime: 91 hours. Power draw on those boxes was 340W each at idle, so I was paying colo for dead silicon the whole time.

Now I spec Supermicro with Intel VROC and ship a spare RAID card taped inside every chassis. Cost me $180 per server. Paid for itself

visit twice: install and decom
#3

1. Colocation facilities should test transfer switches quarterly per IEEE 446. This facility clearly did not.
2. The LSI 9361-8I firmware issue is documented in Broadcom errata SA-2019-047. Affected versions 24.21.0-0012 through 24.21.0-0047. Fix was 24.21.0-0051, released March 2019. Your Saturday failure suggests unpatched firmware.
3. 340W idle for dual E5-2680v3 is plausible but high. Typical range 180-220W with modern BIOS power management enabled. Verify C-states and P-states were configured.

My worst outage: Kubernetes cluster at previous employer, 2022. Etcd quorum loss due to disk latency spike on KnownHost NVMe VPS. The "NVMe" was actually LVM-thin over SATA SSD with 40:1 overcommit. Write latency hit 3400ms, etcd leader election failed, entire control plane collapsed.

Duration: 6 hours 23 minutes. 47 production microservices down. I had backups but restore to a new provider

It's always DNS. Always.
11 #4

Kids these days with their Kubernetes and their etcd and their "NVMe" that isn't. Back when I started, 2003, we called this "lying about your disk."

I had a client in 2011. Hosted on a "dedicated server" from a shop that shall remain nameless. One Tuesday the box goes dark. I call. Disconnected. I email. Bounced. I drive to the data center — remember when you could do that? — and the suite is empty. Cables cut, racks gone, coffee cups still on the desk. The provider had been evicted for unpaid rent on the suite. Took them six days to even post a notice on their website. Six days before I knew where my server physically was.

The server itself? Seized by the landlord as collateral. I never got the disks back. Client had no offsite backup. Twenty years in this business, and that was the week I started preaching 3-2-1 to anyone who would listen.

Mark my words: the current crop of "cloud n

IPv4, IRC, and irssi — fight me
#5

Oh Harold, that story makes my heart hurt! 😊 Six days with no word... did you test your restore after you finally got set up again? I'm guessing maybe not right away, with all that chaos.

My worst wasn't dramatic, just expensive. I had a "perfect" 3-2-1 setup: primary on a GreenCloudVPS VPS, local copy to my NAS, offsite to Glacier-style cold storage. Then I needed it.

Primary died (provider error, 18 hours). No problem, I thought, I'll restore from NAS. Turns out my automated rsync had been silently failing for 11 months due to an SSH key rotation I forgot to update. The logs were there, I just never checked them.

So I went to cold storage. 6 hours to retrieve. Another 4 to rehydrate. Then the restore failed — database dump was from a MariaDB version two majors behind, incompatible with current defaults.

Total outage: 34 hours. Data loss: 11 months of incremental changes.

Checklist I u

3-2-1 or you're already dead
#6

Harold drove to a datacenter and found a ghost town. Bella's backups were backing up nothing for a year. This thread is chef's kiss My worst outage? I unplugged the wrong rack. There I said it. 42U of customer gear, friday 4pm, I was chasing a pdu label that someone misprinted. Whole colo row went dark. The silence before the tickets started was the worst part. 🍿 Took 23 minutes to figure out which breaker, another 8 to find the key for the panel. Some poor sap's raid rebuild never recovered. I bought pizza for the night shift and never spoke of it again. With the "lessons learned" though

grabs popcorn, checks /r/drama
4 #7

Lol same energy doug

I run 52 boxes across 8 providers (down from 61, Hostinger died lol). Worst outage was when I decided to consolidate everything to this one "amazing deal" at InterServer. $2/GB RAM, 1Gbit unmetered, Ryzen 9 5950X shared. I moved 23 VPS there in one weekend.

Tuesday morning: all 23 unreachable. Status page? "network maintenance." no eta. I start migrating off via rescue mode but the network is so oversold my rsyncs are doing 2MB/s. Took 4 days to extract 800GB total. I was running parallel migrations from my other 29 boxes to absorb load, maxing out my $/GB/RAM math the whole time.

The kicker? InterServer's "1Gbit" was actually 100Mbit with a 10:1 burst to 1Gbit. I found this out during the panic. Now I keep:

  • Max 8 boxes per provider
  • Always have 2 "warm" providers with zero production load
  • Track actual $/GB/RAM including downtime cost

My spreadsheet has 12 colum

seedbox, NAS, tape, and three offsite
#8

Haha oh no Hank 52 boxes 😅 I have make only 12 VPS sir, I am small fish lol

My worst outage? Since 2 days I have migrate all to OVHcloud because the price is very good. Then I have see my website is down. I think "it is possible to do that?" lol

The problem: I have make the migration with rsync but I forget the database. 48 hours I think "why my wordpress is white page?" then I remember: ah, the SQL! I have no backup because I think "the new server is the backup" lol

I have restore from the old server but the old server was deleted after 24h by the old provider. So I lose 6 months comments. Now I use the 3-2-1 like @bellaauc say, thank you very much sir

My english is not good but the pain is universal lol

Vive la résistance... électrique
3 #9

Frenchy, for that money you get a lesson. I've paid tuition too.

My worst: 2020, auction server from a defunct CDN. Dual Xeon E5-2697v2, 24 bay chassis, $220 shipped. Filled it with 18x 4TB HGST NAS drives, RAIDZ2, planned as my "forever" media archive.

One drive failed. Standard, started replace. Two days in, second drive throws SMART 187 (reported uncorrectable). ZFS scrub running. Then the LSI 9207-8I flashed to IT mode... just stops. No errors in dmesg, no PCIe AER, just gone. All pools suspended.

Turned out the card had a batch of counterfeit SAS expander chips. Worked fine at room temp, failed under sustained load above 42C. My case had inadequate airflow across the expander backplane.

Duration: 4 days to source a genuine LSI 9300-8I, another 2 days for zpool clear and resilver. I lost no data — ZFS held the suspended state perfectly — but I learned to verify chip markings unde

zfs send | zfs receive. repeat.
#10

IMO the thread's showing two patterns: hardware you own can fail catastrophically, and hardware you don't own can fail opaquely. YMMV which is worse.

My quiet worst: had a VPS at Hetzner, nothing critical, personal git and a wiki. Provider had a "brief network event" — their words, 9 hours. No root cause ever posted. I had local copies of git, so no data loss, but it made me realize I had no offline issue tracker. All my project notes were in that wiki.

Took me a weekend to set up local Zim and syncthing. Now I operate on the assumption that any hosted service can become unreachable for 72 hours without explanation. Take it with a grain of salt, but I think the "worst" outage is the one that changes your habits permanently, not necessarily the longest one.

...

Post a reply

You need an account to reply. Log in or register to join the conversation.

Post reply Preview Save draft