Skip to content

Hetzner: 6 months fine, then 2 weeks of daily outages

Reviews by harbourops 6 replies 472 views
#1

I had six months fine, then 2 weeks of daily outages.

  • 99.7% uptime for my account over 6-month baseline
  • 14 distinct outage events within 14-day window beginning August 1
  • All events: under 45 minutes duration
  • Root cause: under investigation

Network specifications during incident window:

  • Upstream: dual blend, Cogent + HE
  • No planned maintenance logged
  • DDoS mitigation: automatic, no trigger events

I do not confirm or deny datacenter-specific hypotheses. I keep about thirty client sites running in Amsterdam. (see their status page: https://status.hetzner.com)

#2

Re-upped my re-newal at Hetzner pre-outage, now re-gretting. Domain not domain name, just domain. WHOIS shows their IP re-delegated from AS pre-viously stable, now backordered into different netblock. Dropcatch mood: activated. Their re-seller status re-veals nothing. Pre-emptive transfer initiated. Domain not domain name.

#3

V4 path changed. V6 stable. BGP prepended 3x on v4 only. Fw logs: no drops. Rt trace: latency spike 180ms to 340ms, same hop every time. Peering? No. transit insult: cogent. Tier-1 compliment: none found. Cfg check: normal. BGP community 2914:666 present = "performance issue upstream" but no ack from Hetzner.

#4
ned69 said:
BGP community 2914:666 present = "performance issue upstream"

The BGP community 2914:666 you observed is Cogent's well-known community for degraded path, not a performance issue per se but a traffic engineering signal. What's more significant: Hetzner's IPv4 prefix originated from AS64201 during the incident window rather than their usual AS398721. This suggests either prefix hijack, intentional anycast shift, or datacenter-level traffic redirection. Their RPKI ROA covers /22, but the more-specific /24 was invalid. IRR route objects in RADB and RIPE both stale since March. No IRR route6 objects at all for their v6 space, though RPKI valid. The v4/v6 asymmetry you note is consistent with selective datacenter evacuation, not upstream failure. I'd want to see their BGP announcement history before and after.

iBGP, eBGP, don't care, just peer
#5

What could go wrong with datacenter evacuation? EVERYTHING. Your data crosses JURISDICTIONS without consent. GDPR violation? Maybe. WARNINGS: fail2ban rules don't migrate, firewall state LOST, new IP reputation UNKNOWN, previous blocklists CLEARED. Attack surface EXPANDS during transition. Document every packet! — https://status.hetzner.com

airgapped, encrypted, faraday'd, still worried
5 #6

Six months fine then two weeks daily outages thats the pattern that gets you thinking (and I do mean thinking deeply about this) because its not random its structured its like someone flipped a switch (or more accurately like someone didnt flip a switch when they should have) and the timing august heat maybe datacenter cooling (I worked in one once it was terrible the noise the noise alone) but nobody confirms anything and these hosts never confirm they just deflect with statistics (99.7% sounds good until youre in the 0.3%) and I keep seeing this pattern across providers not just Hetzner (though Hetzner is the one were discussing here obviously) where stability breeds complacency complacency breeds corner cutting and then suddenly youre explaining to your client why their ecommerce site was down during lunch rush again and the host says "under 45 minutes" like thats acceptable lik

#7

What upstream blend before august 1, same cogent+he or different

Post a reply

You need an account to reply. Log in or register to join the conversation.

Post reply Preview Save draft