Skip to content

Surviving your first provider death spiral

General Discussion by tamuma 25 replies 2.1K views
5 #1

My first provider just went dark for 72 hours with nothing but "we're upgrading" on the status page

I'm still pulling backups through a rescue ISO

What signs did I miss, and how do I get out alive next time

#2

Pro tip: this checklist saved my free tier migration twice.

  • Check their status page history. More than 3 "unplanned" incidents in 90 days is a pattern.
  • Search their announcement forum for "we're upgrading" with no completion date. Heads up: this usually means hardware failure they can't afford to fix.
  • Test their support response time before you need it. Open a billing question. If it takes 48h+, imagine downtime.

I documented my escape from RackNerd to Oracle Cloud here. The 200GB always-free block storage is underrated for this.

licensing is a suggestion
#3

My clients depend on my infrastructure choices. I only use premium providers with premium support. If there are red flags I open a support ticket immediately. The premium hosts respect that.

My clients have never lost data. Premium SLA guarantees that.

IPv4, IRC, and irssi — fight me
#4

Actually (back in my day (which was 2019 (specifically running Proxmox VE 5.4-3)))) we had a different term for this. "Stability event." The provider (won't name names (actually it was a small operation in Germany running OpenVZ (not even KVM!))) sent "we're upgrading" and then... silence. For six weeks. I had migrated everything to a self-hosted XCP-ng 8.0 setup by then. Nested virtualization was a nightmare but you do what you must.

#5
tamuma said:
My first provider just went dark for 72 hours with nothing but "we're upgrading" on the status page I'm still pulling backups through a...
  • The "we're upgrading" email is the kiss of death
  • Especially if it mentions "improved experience" without specs
  • Sub-item a: check if their WHOIS privacy expired
    Sub-item b: that means they didn't auto-renew
  • Happened to me with a "premium" reseller
4 #6

Did you test your restore? 😊

The 3-2-1 rule isn't just for backups, it's for providers too. 3 copies of your data, 2 different hosts, 1 offline or air-gapped.

Before I migrate I always:

  • Spin up identical stack on destination
  • Restore from backup
  • Verify checksums
  • Keep old host running 48h

The checklist is good but the real test is "can you actually leave in 24 hours?" Most people can't.

3-2-1 or you're already dead
#7

Which RackNerd plan had the 200GB

#8
hoshinobro said:
Which RackNerd plan had the 200GB

That was the 2.5GB KVM Black Friday 2022 deal. 200GB SSD, 5TB transfer. Gone now. I only mentioned it because the migration path matters more than the plan.

Oracle Cloud's always-free tier is what I actually run on now. 200GB block storage per account, two ARM instances, no expiration.

licensing is a suggestion
#9
bellaauc said:
The 3-2-1 rule isn't just for backups, it's for providers too.

This is cargo-cult resilience. Running five $5 VPSes doesn't give you five nines. It gives you five times the alert fatigue and five times the attack surface.

My clients get RAID-10 on NVMe, offsite replication to a second premium facility, and a support contract with a human who answers in under 15 minutes. You can't checklist your way out of cheap infrastructure.

IPv4, IRC, and irssi — fight me
#10
haroldgsm said:
This is cargo-cult resilience.

My five providers don't alert because they're not running anything that needs 99.999%. They're warm spares. The whole stack restores from Terraform in 40 minutes.

Your premium SLA didn't help when OVH Strasbourg burned. Premium just means premium price. The fire didn't check contracts.

3-2-1 or you're already dead

Post a reply

You need an account to reply. Log in or register to join the conversation.

Post reply Preview Save draft