Skip to content

Monitoring alerts are just anxiety generators

General Discussion by Joy 7 replies 535 views
#1

I am run 12 VPS for plex and xixixi some client sites... pefore I am have PagerDuty and very much alert... CPU 80%? PING... disk 90%? PING... every night I am wake up for nothing... xixixi

3 month ago I am turn OFF all alert... now I am check log every sunday only... guess what? Same problem count... but I am sleep better... my MTTR actually DOWN because I am fix real problem with clear head... not panic at 3am for spike lah

I am think alert culture is make us reactive zombie... not engineer... xixixi

What you all think... am I crazy or genius lah

#2

PagerDuty is just expensive panic alarm lah

#3

1. I understand you very much
2. I have same with Nagios before... very noise... no sleep
3. Now I use simple script check weekly... much better
4. But I keep one alert: disk full... that one is save me twice
5. Balance is key... not all or nothing... thank you

You are brave for say this... many people are angry

#4

Sorry for the silly question but... if you turn off alerts how do you know when the site is actually down? I run a small shop and my customers email me when its broken... feels like I should be first to know?

I tried UptimeRobot free tier but got so many false alarms I stopped checking... maybe weekly logs is better for someone like me too?

frames, tables, still valid HTML
7 #5

Up, up, up

Your MTTR is down because your problems are small... wait until /24 will cost you a kidney and you miss a BGP hijack for 6 hours because "sunday log check"... alerts have a cost... silence has a bigger one

#6
Joy said:
My MTTR actually DOWN

Anecdote vs data. Virtualization tax on your attention is real, but so is the cgroup oom-kill you sleep through.

I ran OpenVZ nodes for years without proper alerting. "Stable" until host kernel panic at 2am and 40 containers gone. Took 4 hours to notice. With KVM + proper kernel panic alerts? 8 minutes.

Your 12 VPS are toy workloads. Scale past single-node and this breaks.

virsh list --all | wc -l: 47
#7

Same with my Nagios alerts, just noise

#8

Which kernel panic handler, netconsole or kdump

Post a reply

You need an account to reply. Log in or register to join the conversation.

Post reply Preview Save draft