Raid 6 on a Contabo box | 8x 16TB | expand from 64 to 80TB | mdadm --grow /dev/md0 --raid-devices=9 | background reshape starts | I go to sleep | grep disaster /var/log/kern | 3AM | all 8 original disks dropped simultaneously | kernel spat out thousands of ata errors | array gone | production vm storage gone | cat backups | grep recent | wc -l | zero | the ninth disk the new one it kept responding fine | firmware bug in lsi 9305-16I | expand triggers timeout cascade | vendor confirmed | cve pending | check your expand schedules | sort -u | panic
RAID failures: lost the array during a routine expand
This is exactly why I run two separate docker compose stacks across different machines with zfs send/receive to my basement box behind a reverse proxy | your autopsy saved me man I was about to expand a similar array this weekend | switched to mirrored vdevs years ago for my own stuff but the dayjob still has legacy raid | might finally convince them
Back in 2009 I lost a 3ware 9650 array to a firmware bug during rebuild | they don't make them like that anymore except apparently they absolutely do | the smell of hot sas backplanes at 3am is eternal | my condolences on the vm storage | did you have offsite or was this the classic single-location prayer
Dear sir/madam pavel_train, kindly confirm whether the replacement firmware is available only or we must await the cve publication itself | my own expand is scheduled for next week at InterServer only | I am running the 9305-16I card itself | this is most alarming news | kindly share the vendor ticket number if possible
Lsi 9305-16I firmware version before the expand?
Ninth disk kept responding fine makes no sense