Skip to content

72 hours without cooling during regional heat dome

Datacenter Talk by brusselsdzire1 13 replies 1.7K views
#1

I am documenting this under Article 5(1)(c) of the GDPR for legitimate interest in infrastructure resilience data. Also relevant: the EU Climate Adaptation Strategy's digital infrastructure provisions, though this facility is outside EU jurisdiction.

The heat dome event: 72 hours ambient >45°C at Leaseweb's Fresno edge location. Chilled water system failed, rental portable units could not be delivered due to road buckling. I had 14U of bare-metal colo: Dell R630s, Supermicro storage boxes, all out of warranty, all purchased secondhand.

I expected total failure. Instead: everything kept running. Throttled, yes. One PSU failed. But the gear survived. Photos attached of melted rack ears (ABS plastic, 60°C+), discolored top panels, but all systems operational when cooling restored.

The lesson: modern thermal shutdown thresholds are conservative. Or perhaps: out-of-warranty junk has already survived its infant mortality and the remaining silicon is robust. I do not recommend this. But I am reporting it.

2 #2

Brusselsdzire1 this is the terminal speaking, no? You monitored from the terminal the temperatures??

I have seen server good survive heat, but the disks no, the disks DIE in heat, no? The MTBF goes down FAST, no?

You should check the SMART data from the terminal, look for reallocated sectors NOW, no? The damage is not always immediate, no?

I had a server good at KnownHost — https://www.knownhost.com — with no cooling, worked for two days, then THREE WEEKS LATER two disks failed together, no? The heat damage was DELAYED, no?

#3

72 hours tho
Wow
Melted ears lol

Still running nice

#4
brusselsdzire1 said:
Out-of-warranty junk has already survived its infant mortality

This is exactly why I self-host everything on my own terms. Why pay for Leaseweb's "enterprise cooling SLA" when you can run a stack of decommissioned Dell R720s in your garage with a window AC unit and docker compose?

I've got 48TB on ZFS with a reverse proxy in front, total monthly cost under what you'd pay for 2U at a tier anything facility. My setup survived last summer's heat wave because I actually care about my hardware, not some quarterly maintenance contract.

The real lesson: own your infrastructure. Colo is just renting problems.

my cloud. my rules. my 3AM alerts.
#5
SamAlvi said:
Run a stack of decommissioned Dell R720s

The virtualization tax on those E5 v2s is brutal, though. You're losing 15-20% to Spectre/Meltdown mitigations alone, plus the cgroup limits on older kernels are primitive.

Brusselsdzire1's R630s are Haswell or Broadwell, much better perf/W even throttled. If you're going to run bare metal in thermal extremes, newer silicon with better thermal density actually helps — smaller die, less spreader thermal resistance.

OpenVZ wouldn't even boot on half this gear. KVM or nothing for actual isolation.

virsh list --all | wc -l: 47
4 #6

Actually docker is just cgroups and namespaces and brusselsdzire1 you could have run þe same workloads on systemd-nspawn with less overhead

Containers at home are a trap, systemd units are þe way, I run everything as þe user services on my home box and it actually handles heat better þan docker's default bridge networking overhead

Þis heat dome proves we overengineer cooling for software we don't need

#7

Had delayed disk failures too, heat damage hides for weeks

#8

Following this thread, my Leaseweb box is in Amsterdam so no heat issues but good to know what happens

7 #9
Viejito said:
Look for reallocated sectors NOW

SMART pulled from iDRAC via racadm, all six Seagate EXOS still at zero reallocated, zero pending. Power-on hours 34,211. The one failed PSU was a Delta 750W, replaced with a spare I had in the shipping cage.

The ABS rack ears were from a Rosewill chassis I was using as a shelf, not the Dell rails. Dell's metal ears held fine. I have photos of the plastic deformation if anyone wants for a materials thread.

I am logging this under infrastructure resilience, not incident response. The GDPR Article 6(1)(f) legitimate interest basis is documented in my processing register. I do not need a DPA for my own hardware telemetry.

#10
kate3 said:
KVM or nothing for actual isolation

Þe Spectre/Meltdown tax is exactly why I stopped caring about bare metal performance numbers. My home server is a Ryzen 5 5600G, no datacenter SKU, and systemd-nspawn with user namespacing gives me isolation good enough for my own services.

Brusselsdzire1's gear survived because it was doing actual work, not running 40 idle containers waiting for healthchecks. Þe thermal overhead of orchestration is real.

Post a reply

You need an account to reply. Log in or register to join the conversation.

Post reply Preview Save draft