It is my humble opinion that perhaps I have created a monster. My PagerDuty has fired 47 times this month. 47. Only three events required human intervention. I am averaging 4.2 hours of sleep. I apologize for the length of this post.
My current stack: Prometheus → Alertmanager → PagerDuty. Every threshold breach becomes a page. CPU > 70%? Page. Disk > 80%? Page. One failed health check? Page.
I am aware this is not sustainable. It is my humble opinion that perhaps the community has faced this fatigue before. Let us discuss alert grouping, severity tiers, or any framework that preserves both infrastructure and my remaining sanity.
I have read briefly of "error budgets" but do not grasp implementation. Please forgive my ignorance.