Skip to content

Mini-guide: reading between lines of provider status pages

General Discussion by beto60 8 replies 906 views
#1

I am reading the status pages for three years now, caramba! The providers think we are not seeing the patterns, nossa! But we see.

"Scheduled maintenance" at 3AM on sunday = they are not having redundancy, kkkk. "Degraded performance" for six hours = someone kicked the cable and they don't know where.

I am making this list for the community. What else you are seeing?

#2

Lol the classic "we are investigating" for 45 minutes then "all clear" with zero explanation means they rebooted it and prayed "intermittent connectivity" = the intern is learning on production I had a provider say "enhanced security measures" once that was them locking themselves out of their own rack The best is "third party network issue" when its always the same third party and you check and its literally their only upstream

2 #3
beto60 said:
They are not having redundancy

This is exactly why I self-host my status page. Uptime Kuma in docker compose, reverse proxy through nginx, I get actual data. Why pay for someone else's "all systems operational" theater when you can self-host real monitoring.

I have seen "scheduled maintenance" translate to "we never patched this and now it is forcing us to reboot." The providers who post actual maintenance windows with RFC numbers and change control? Those are rare. I bookmark them.

My favorite decoding: "increased error rates" means their database replica is lagging and they do not want to say "replication." I have the graphs to prove it from my own self-hosted stack.

my cloud. my rules. my 3AM alerts.
#4

"Routing optimization" = they switched upstreams to save money

Route via Cogent adds 12ms. Route via HE adds 8ms. I have traceroutes from six providers showing this pattern.

Provider A (2024-03): avg 34ms via Level3
Provider A (2024-06): avg 52ms via Cogent
Status page: "network improvements"

The numbers do not lie. The words do.

1ms or I don't want it
#5

"brief interruption"...... five hours...... "minimal impact"...... everyone in the thread screaming...... "we appreciate your patience"...... they never appreciate the patience...... just the silence...... especially in lists...... "affected services"...... then they list everything...... except the thing that actually broke...... you get the idea......

#6

"we are aware" = please, server is on fire and they just woke up

"escalated to senior team" = no one is home, "rerouting shortly" = they are googling the error message right now

No articles needed, we see through it

#7

U kids with your statuspage.io templates and "incident retrospectives" that say nothing.

I saw a provider write "root cause: unexpected system behavior" :-) thats called a bug, friends. We used to say bug. "unexpected system behavior" is what you say when you r scared of lawyers.

Someone should make a bingo card. "enhanced monitoring" = we bought pingdom. "we have learned from this" = we will do it again in 6 months.

#8

Unexpected system behavior" is just bug with a lawyer hat

world record: 4min Arch install
#9

The pingdom one is too real lah, I see this everywhere

Post a reply

You need an account to reply. Log in or register to join the conversation.

Post reply Preview Save draft