Monitoring 10+ cheap boxes without going crazy
This thread drifted fast. Back to MeritBudi: you have alert fatigue because you alert on symptoms, not on business impact. Certificate expiry? User complaint is the alert. Disk full? That's a symptom, the impact is service down. Define SLOs. One external probe for "can user reach my service." One internal probe for "will it stay up." Everything else is dashboard.
In Latvia we cannot afford 14 VPS. I have two, on Hetzner. One runs monitoring, one runs everything else. If monitoring dies, I know because the other one is also down. This is not sophisticated but it is cheap.
Two-node monitoring with no redundancy is how you get paged at 3am because the monitoring host OOM'd. I did this in 2008 with Nagios. Never again. Minimum three nodes for consensus, or accept that you will have blind spots.
Following.
Same here. I have 8 VPS and Uptime Kuma on a Raspberry Pi at home. The Pi is the single point of failure but it's been up 300 days so I stopped worrying.
300 days uptime on a Raspberry Pi means you haven't patched kernel in 300 days. I hope it's not exposed.
It's behind NAT, relax.
Ok ok I learn already. So maybe I do:
- One Uptime Kuma external for blackbox
- One Netdata or VictoriaMetrics internal for metric
- Only 4 alert rule like uma say
But which provider to put the monitoring? If I put on same provider as my main site, when provider down I get no alert also.