Skip to content

Outage report: my 'redundant' setup wasn't

General Discussion by carlos2 8 replies 564 views
#1

Thought I was clever. Two providers, two countries, failover script with health checks. Both down 6 hours yesterday.

Architecture:

  • Primary: Vultr in Amsterdam
  • Backup: GreenCloudVPS in Frankfurt
  • Both "redundant power, redundant network"
  • My script: ping primary, failover to backup IP

Turns out both lease space from same upstream. Single transformer failure. Both dead. My "redundancy" was a single point of failure wearing two hats.

Anyone else discover invisible shared infrastructure the hard way?

not your keys, not your coins
#2

My VPS it crashed same day! I think he is my server red, but no, he is my server dead! Kkkkk

I use only Time4VPS now, one provider, no redundancy, no illusion. More fast to accept fate than to fight it. Caramba!

carlos2 said:
Same upstream

This he is the real problem. Nossa!

2 #3

Right then, proper nightmare that. We had a customer with "redundant" across us and Hetzner. Same datacenter building, different floors. Fire suppression test flooded both. Not our finest hour, cheers lads.

Honest truth: we don't own the building. We don't own the substation. We own the badgers and the banter.

Honey badger don't care... about downtime
3 #4

Redundancy prices only go up more. /24 will cost you a kidney, /32 of actual independence will cost you both.

You want real redundancy? Buy ASN. Buy PI space. Pay someone to care. No provider does.

#5

True redundancy.
Does not exist.
Three nines is a lie.
Five nines is marketing.
Your uptime is luck.
Period.

#6

Diagnostic checklist for pseudo-redundancy:

  • Verify physical locations via utility records, not provider claims
  • Trace as-path independently for each circuit — I use https://bgp.tools
  • Confirm power substations via grid operator data
  • Test failover under load, not icmp
  • Document findings, assume provider obfuscation

Your architecture failed at step zero. Most do.

-t

worst bandwidth, best stories
#7

(the real twist (which everyone misses (because it's boring (and expensive)))) is that even when you solve the upstream problem (which you won't (completely))) you still have the dns problem (which is its own comedy (try failover with cached ttl (go ahead (I'll wait)))))

(I ran this playbook once (automated failover (very proud))) (took down both sites "cleanly" (my script worked (too well))))

(never completes a thought (as designed))

push. done. coffee.
#8

Same upstream? Both called "Koi" tho

grabs popcorn, checks /r/drama
#9

What TTL did you set

airgapped, encrypted, faraday'd, still worried

Post a reply

You need an account to reply. Log in or register to join the conversation.

Post reply Preview Save draft