Skip to content

Monitoring: my alert fired, I ignored it, consequences

General Discussion by nate_pad 5 replies 139 views
#1

My alert fired at 3am for disk usage on my vps, I was so tired and I thought it was another false positive so I just muted it. Then this morning my site was down and my host said it was a cascade failure starting from that server?

Is this normal? Should I have not ignored it? Sorry if dumb question

4 #2

Just use Prometheus with proper alertmanager routing. Skill issue.

The real problem is you're running a VPS without understanding what the alerts MEAN. CAPS for emphasis: DISK FULL CAN KILL YOUR DATABASE.

oops: 0000 [#1] SMP
#3

Much noise people tune out.

I once had a RAID controller scream for three days before it actually failed. Ignored it because "it's always complaining." Lost everything. The alerts knew before I did.

SPARCstation 20, still serving HTTP
12 #4

Today I learned that ignoring alerts is bad forever apparently I do this all the time my pager goes off like 20 times a night so I just sleep through it now maybe I should fix that huh

grabs popcorn, checks /r/drama
#5

Alert fatigue is real. The industry knows this. Pagerduty wrote papers. Nobody implements the fixes. Your brain tunes out. The signal drowns. Then cascade hits;

your margin is my opportunity
#6

¡¿is very common, amigo?! The v/b confusion in monitoring is real, "vital" vs "bital" metrics jajaja

I write long post then delete: shorter version. You need alert severity. Page only for pageable. Everything else email. I learned hard way.

Post a reply

You need an account to reply. Log in or register to join the conversation.

Post reply Preview Save draft