Skip to content

OVH fire 3-year anniversary, what changed?

VPS Hosting by LichunLars 25 replies 2.7K views
5 #1

I ran traceroutes to every provider I could find still operating from the same metro areas. Results are mixed.

[code]Contabo Munich: 14.2ms avg, route via 162.254.x.x
OVHcloud Gravelines: 11.8ms avg, route via 51.178.x.x
Vultr Frankfurt: 22.4ms avg, route via 198.244.x.x (detour!)
Hetzner Falkenstein: 13.1ms avg, route via 141.94.x.x

Route via Vultr's Frankfurt adds 12ms compared to their old path. They moved, or their upstream did. I checked https://status.vultr.com but no maintenance posted.

Physical redundancy improved. Logical redundancy? I see more "availability zone" marketing, less actual diversity. Same building, different VLAN, call it resilient.

What are others seeing? Concrete numbers, not press releases.

1ms or I don't want it
#2

Cat /dev/disaster | grep "fire suppression" | sort -u | zgrep "annual report" | awk '{print $2}' | sed 's/compliance/actual water damage/' | the more things change | pavel_train@kovalevskaya | ssh node | uptime | 99.9% | bc | 0.1% downtime | 8.76 hours | wc -l | one provider learned | another did not

#3

Ngl the real change is everyone moved their LLM inference boxes to "geographically diverse" setups lol

I had 4x A100s at the old site, now I split 2+2 across InterServer and Time4VPS. Training run gets checkpointed every hour, resume from either side. The VRAM math works out same, network overhead is whatever, but I sleep better

Same risks elsewhere tho, I looked at KnownHost and their "two locations" are same metro, same power grid, probably same flood zone. Lol

CUDA cores are my love language
#4

Scrap gods gonna scrap god,,,,, ,,,,, Three years and we still chasing the same cheap dedis from the same few sellers,,,,, I remember when OVHcloud had that flash sale the day after the fire,,,,, Everyone panic moved and then the new place had a power issue six months later,,,,, The scrap gods dont care about your uptime,,,,, Community funded sre for bytearray actually worked though,,,,, We got a discord channel with a guy who monitors their actual racks with ipmi,,,,, Costs us fifty bucks a month split twenty ways,,,,, Best insurance I ever bought,,,,, ~~dave

#5

1. Actual improvements I have observed:
A) More providers publish physical location data
B) Some offer cross-site replication by default
2. Same risks persist:
3. Single metro "redundancy"
4. Shared power infrastructure not disclosed
5. Community SRE model:
I. Works for small providers
Ii. Does not scale
Iii. Still worth doing

#6

Same. Had "diverse" racks, same substation. Both went dark.

#7

Which providers actually publish substation-level data? I haven't seen it.

#8

Which metro? InterServer and Time4VPS don't list that.

$3/year. 128MB RAM. Pure happiness.
#9

Larryjeong Secaucus and Vilnius yeah. Latency between them is trash for synchronous stuff but my checkpoints are async. 2+2 A100s at each, training resumes from whichever has the latest. Time4VPS is just CPU head nodes with storage, the actual compute is colo I found through a contact. Not naming the facility because their fire suppression is probably a guy with a bucket.

CUDA cores are my love language
#10

Garykwh how are you getting A100 colo for less than cloud? Every quote I see in Frankfurt starts at 800 EUR per U.

mitigated 800Gbps before breakfast

Post a reply

You need an account to reply. Log in or register to join the conversation.

Post reply Preview Save draft