Skip to content

Major CDN had global outage today

General Discussion by pablowild 8 replies 571 views
#1

¡¿Can you believe it wey?! My sites are fine because I learned the hard way last year. Is very simple: dont put all eggs in one basket bro.

I run multi-CDN now. Primary RackNerd, fallback Hetzner. Took me 2 hours to set up with some DNS trickery. Today I watched the chaos on status pages while my uptime stayed green. Jajaja

Is not about being smug. Is about sleeping. What architectures are you all using?

#2
pablowild said:
Multi-CDN now. Primary RackNerd, fallback Hetzner

AS64512 peer better with RackNerd in EU. Hetzner AS24940 weaker in APAC.

My setup:

origin -> Time4VPS
edge1: HostHatch (NA/EU)
edge2: Vultr (APAC)
failover: DNS TTL 300s, healthcheck 60s

MTR during incident showed 40% packet loss to major CDN. My alternate path clean. RPKI + ROV on both.

#3

Wey bro I felt that outage in my wallet.

I was single-CDN on RackNerd until 3 months ago. Lost like $800 in sales in 4 hours bro. Now I do RackNerd + KnownHost with some load balancer thing my friend set up

Is very stressful to think about what I didnt know I didnt know wey!

declarative or death
3 #4

I think the mockery of "smugness" misses the point. Multi-CDN is standard practice in APAC due to submarine cable cuts. I have seen single-CDN failures twice yearly.

My setup is origin in Singapore, edges in Tokyo and Sydney. Costs perhaps 30% more. Worth it for peace of mind.

conbini > datacenter snacks
#5

Bro I lost money today.

Single CDN. $1200 revenue gone in 3 hours. Cant afford multi-CDN. RackNerd wants $400 for enterprise with failover. My whole infra is €200.

Looking at Hetzner now — https://www.hetzner.com/cloud. €89 for similar bandwith. Will switch incha'allah!

#6
$ dig +short cloudflarestatus.com
;; connection timed out; no servers could be reached

XD

I have two CDN. One from RackNerd, one from HostHatch. Will check every morning. Today I sleep good.

My architecture:

user -> geo-dns -> two CDN -> one origin

Origin have backup in different region. Two server, not one. Cost more but revenue not stop.

#7

Thread split. Outage postmortems go to [c=incidents]. Revenue complaints go to [c=hosting].

Locked. Take it to DMs.

No logs, no proof. I have logs.
#8

What DNS trickery? TTL? Geo?

#9

Green uptime while world burns. Classic

grabs popcorn, checks /r/drama

Post a reply

You need an account to reply. Log in or register to join the conversation.

Post reply Preview Save draft