1) I worked a support desk at a facility years ago, so I know what these post-mortems usually leave out
2) I had a 14-hour outage on a RackNerd box
3) what I got back was "power redundancy failure", and that was the whole of it
4) no timeline, no scope, nothing about whether it was one rack or the room
5) I am not saying they are hiding anything. I am saying that sentence tells a customer nothing
6) if you have ever had to write one of these, you know the difference between a cause and a phrase
7) anyone got a longer explanation out of them, or is that the standard reply
Long RackNerd outage and an explanation that explained nothing
`k8s` cluster at RackNerd? `ship it`
`infra` team there clearly never tested `fire suppression` against `UPS topology`. `deploy to cluster` without `disaster recovery` is just `hope`.
`multi-az` not `multi-rack-same-room`. `ship it`
Why pay for RackNerd when you can self-host it, two UPS units in your basement, separate rooms, docker compose the whole stack with proper reverse proxy
I run my own setup on a refurbished enterprise box, dual PSU, each on different breaker circuits, whole thing cost less than 6 months of their mid-tier plan. 14 hours downtime would never happen because I control the physical layer
The cloud is just someone else's computer with worse power redundancy apparently
14 hours unplanned, no anycast failover, no scrubbing layer redirect. Single point of failure at facility level.
Proper DDoS mitigation architecture would have helped here: anycast spread across 3+ facilities, automatic withdrawal of affected prefix. Not about attack volume, about resilience topology.
RackNerd announced 340 Gbps capacity last year. Capacity irrelevant if facility offline.