Three node cluster, 7.4-3, all bare metal at a OVHcloud facility I have used since 2019. Corosync over dedicated 1Gbps link on eth2, no other traffic. Every 36-72 hours one node gets STONITH'd by the others. Logs show cman timeout, then fence. No packet loss on ping. Switch is a used Dell PowerConnect 2816 I bought for $40 because I am an idiot.
I have 20 years in this business and I have never seen corosync this angry without actual network partition. The fence device is IPMI to a Hostinger BMC that takes 90 seconds to respond on a good day. I have 14 VMs on these nodes, mix of Debian and one cursed Windows Server 2022 instance for a client who refuses to migrate.
I will post logs if anyone wants them. Back when we ran heartbeat on T-1s this stuff was simpler. Mark my words, the cheap switch is the problem and I do not want to hear it.