Skip to content

The outage that wasn't: My monitoring was the broken part

General Discussion by jihoonred86 4 replies 117 views
#1

Woke up to 47 Nagios alerts screaming that my Contabo box was down. Fired off an angry ticket, got a calm reply that their end showed 99.99% uptime. Traced it back to my own check_command timeout being too aggressive after I "optimized" it last week. Embarrassing.

Anyone else burned by their own monitoring stack crying wolf? I'm running 6 boxes across 4 providers and this is the third false alarm this month. Starting to think my $3 VPS with better YABS disk scores than my $10 box is the only thing I configured right. 1874 MB/s on that bad boy.

What's your alert fatigue story?

#2
jihoonred86 said:
47 Nagios alerts screaming

We have found that overly sensitive thresholds are usually the culprit when monitoring appears unreliable. In our environment, we implemented a two-strike rule before any alert reaches a human. This reduced our false positive rate by roughly eighty percent without missing genuine incidents.

It is worth reviewing whether your notification chain includes enough context to distinguish between a failed check and actual service degradation. We learned this lesson gradually over several quarters.

Sam

#3

When I see 47 alerts I just delete all

My zabbix is more fast to say lies than to say true things

I no have patience for monitoring that cry wolf every day

I put 5 min timeout now and no more false alarm thank god

But my ping from brazil to RackNerd is 180ms what u expect for 3 dollar

Vive la résistance... électrique
#4

I made VPS last month and monitoring is very confusing for me sir. I use simple uptime checker from GreenCloudVPS panel, no false alarm yet. Maybe because it is basic 🙏

Your story teach me to check own setup first before blame provider. Very good lesson.

#5

I run 52 boxes so false alarms are my life 😂

My stack:

  • 3 on Contabo (€2/GB)
  • 8 on Hetzner (€1.80/GB, best deal)
  • 12 on Vultr (free tier, "oracle" style)
  • Rest scattered

Alert fatigue got so bad I built a dead man's switch. If I don't ack in 4 hours, my phone actually rings. Real outages: maybe 2 this year. Fake ones: hundreds.

My fix was simple. Check from 3 locations before it counts. Costs nothing, saves sanity.

seedbox, NAS, tape, and three offsite

Post a reply

You need an account to reply. Log in or register to join the conversation.

Post reply Preview Save draft