The quiet first failure
A member drops out and nobody notices — the array keeps serving, so weeks of degraded running pass without the alarm being read. The second failure is simply the first one anyone hears.
Distributed parity lets a RAID 5 lose any single member and keep serving — which is exactly why most of them arrive here having lost a second one mid-rebuild. Here is what parity can compute back, what it cannot, and why the rebuild itself is the most dangerous hour in the life of the array.
Parity means any single missing member can be computed back from the others. A second failure stops the arithmetic — from there the job is physical: make one of the failed disks readable again, and the maths takes care of the rest.
Why a RAID 5 keeps working with a member missing — and why that grace ends at exactly one.
A RAID 5 stripes data across its members like a RAID 0, but for every stripe it also writes a parity block — the XOR of the others — and rotates that duty around the disks. Lose any one member and every missing block can be recomputed on the fly from the surviving blocks in its stripe. The array carries on: slower, unprotected, but intact.
The capacity arithmetic gives one disk’s worth to parity — four 4TB members make a 12TB volume. The protection arithmetic is stricter. One unknown per stripe can be solved; two cannot. A second failure does not degrade the volume further, it stops it.
Nearly every RAID 5 that reaches this bench got here through that second failure — and usually during the rebuild that followed the first.
Rarely the way the manual imagines.
A member drops out and nobody notices — the array keeps serving, so weeks of degraded running pass without the alarm being read. The second failure is simply the first one anyone hears.
A replacement disk forces hours of sustained reads across every sector of every survivor. Ageing disks that coasted for years fail under exactly that load — or one unreadable sector aborts the whole rebuild.
A disk that dropped out months earlier gets forced back online. Its contents describe an older volume, and the controller happily blends the two eras into one corrupt present.
The disks are healthy but the card that knew the order, stripe size and parity rotation is dead or replaced. The data is all there, unlabelled — and fully recoverable virtually.
The discipline that keeps a recoverable array recoverable.
A stopped rebuild has lost nothing yet. A pushed one grinds failing heads across the only copy of the evidence.
Parity rotation makes order matter even more than on a stripe. It can be derived from the data, but knowing it saves hours.
If it left the array weeks ago it holds an old version of your volume, and forcing it in writes that history into the present.
The better half of the failed pair usually is the recovery. Once it images, the array is one disk down and parity does the rest.
Parity turns the physical recovery of one disk into the logical recovery of everything.
Every member is imaged read-only first, healthy or not. The failed pair is assessed and the stronger candidate repaired — donor heads, board work, firmware — until it images as well. That single repair is normally the whole physical job: with N−1 members readable, every missing block becomes computable again.
Geometry — disk order, stripe size, parity rotation, offset — is derived from the images themselves, and the array assembled virtually, entirely outside any controller. The file system is then rebuilt and verified from the assembled volume. RAID work in Manchester is from £500 +VAT after a free 48-hour diagnostic, with a 50% deposit where a member needs physical repair. See also what to do when a rebuild won’t run and the general playbook for a failed array.
In most cases, yes. One failed member is what parity exists for — the missing blocks are recomputed from the rest. Two failures make it a physical job first: one of the failed disks is repaired to imaging condition, and the arithmetic handles everything else.
Usually. The two rarely fail equally — one is almost always in far better condition than the other. Once the stronger one is repaired and imaged, the array is a one-disk problem again and parity fills the gap. Where both are severely damaged, a partial recovery from what reads is still often worthwhile.
Not if the survivors are old, noisy or already logging errors. A rebuild reads every sector of every member for hours — precisely the load that kills a second disk. Image the set first; after that, a failed rebuild costs you nothing.
Because a rebuild is the hardest day’s work the survivors have ever done, taken on when they are oldest. Any weak member tends to fail under it — and even without a second death, a single unreadable sector on a survivor can abort the rebuild on many controllers.
Every member, including both failed ones. Parity reconstruction needs N−1 readable disks, and the best candidate for that final slot is usually the better half of the pair that failed.
From £500 plus VAT after a free 48-hour diagnostic — the same banding as all RAID work here. A 50% deposit applies where a member needs clean-air or board-level repair, with the balance due only on success.