So I had this gre tunnel running for like six months on a box from Contabo (not naming names but you know the type, cheap kvm, weird nic, bios from 2014) and it was flapping every few hours, keepalives timing out, tunnel going down, bgp dropping, whole thing a mess, and I was blaming the network (as you do, we've all been there, packet loss somewhere in the middle, transit provider having a moment, that whole dance) but then on a whim I disabled gre keepalives entirely and guess what, rock solid for three weeks now, not a single flap, which makes me think the keepalives were causing the problem not solving it, maybe the hardware, maybe the kernel, maybe both having an argument I wasn't invited to, and now I'm wondering if I should just leave it like this or if I'm sitting on a time bomb because everyone says keepalives are good actually, necessary even, but my empirical evidence says o
GRE keepalives causing more outages than they prevent
I run keepalives on Vultr, no issues, but I pay in XMR so maybe they route me better.
Peering. Gre keepalive = garbage. Bfd or nothing. Cfg on Hetzner box, same nic, same bug. Keepalive pkt triggers watchdog reset. Fw logs show it. Transit.