A mains failure took the building down. On restart the PowerEdge reported two drives missing from a five-member RAID 5 — which sounds like the end of an array and was in fact two dead circuit boards in front of two perfectly healthy disks.
← All case files · from £500 + VAT
A mains failure took the building down. On restart the server came up but the controller reported two drives missing from a five-member RAID 5. The client checked through the array management utility, confirmed what it was telling them, and tried the controller’s own recovery option, which did not resolve it.
On the array were the accounting system, the customer database and everything operational. They stopped there rather than continuing to try options, and that decision is the reason this recovered as completely as it did.
When a RAID 5 loses two members weeks apart, the first is genuinely dead and the second follows under rebuild stress. When both drop in the same instant, something outside the disks did it — and the disks themselves are frequently fine.
Both ejected members had board-level damage. The mechanisms were sound, the platters unmarked, and the data entirely intact behind a circuit that had stopped answering.
Its recovery routine treats two absent members on a single-parity array as unrecoverable, which for its purposes is correct — it has no way to reconstruct two unknowns from one parity set. That is a limit of the controller, not of the data.
No replacements inserted, no member forced online, no array recreated. The original member metadata survived intact, which made reconstruction straightforward rather than archaeological.
The distinction matters commercially as well as technically. An array that lost disks gradually usually means real mechanical damage and a partial recovery. An array that lost several at one instant from one event is often board-level across the set — more members to repair, and far better prospects.
Both damaged boards had failed in the same place, which is common and is worth understanding because it explains why the disks survived.
A drive’s circuit board carries transient voltage suppression on its power rails — components whose entire job is to conduct a surge to ground rather than let it reach anything else. They are designed to fail closed. When one does, it short-circuits the rail it protects, the drive stops drawing power, and everything downstream of it — the motor controller, the preamplifier, the head assembly — is spared.
So a drive that has been surged is very often a drive whose protection worked. It presents as completely dead, which is alarming, and the mechanism behind it has never been touched.
The instinct with a dead board is to fit one from an identical drive. On anything made in roughly the last twenty years that does not work, and the reason is specific.
Each drive is calibrated individually at manufacture. The head assembly in a particular unit has its own preamplifier bias values, fly-height compensation, per-head read channel settings and a defect list mapping out sectors that were already bad when it left the factory. Those values are unique to that one drive and are stored in a small flash chip on its circuit board.
Fit a board from a physically identical drive and the controller drives this head stack using that drive’s calibration. The result is a drive that spins, sounds healthy and reads nothing — or worse, reads intermittently and inconsistently, which is harder to diagnose than a clean failure.
So both boards were repaired at component level: the failed suppression components replaced, the damaged rails verified, and the original adaptive ROM left in place on its own board. Where a chip cannot be saved it is desoldered and transplanted onto a donor board, which is the same principle by a longer route.
With all five members responding, each was imaged read-only. No mechanical faults surfaced on any of them, and the two repaired members imaged as cleanly as the three that had never dropped.
The array was then reassembled from the images, entirely outside the PowerEdge and its controller. The geometry was derived from the data rather than trusted from metadata: candidate combinations of stripe size, member order, parity rotation and data offset were applied to the images and each tested against whether recognisable file-system structures resolved at the offsets they should occupy. A wrong geometry does not yield a slightly wrong volume — it yields noise — so the arrangement that produces a coherent NTFS structure across several widely separated regions is the right one.
From the assembled volume the file system was extracted, and the databases checked page by page for structural consistency rather than assumed good because they attached.
Everything. Accounting system, customer database and operational file shares, nine days from arrival — most of that being component-level repair on two boards and imaging five drives before reconstruction could begin.
A UPS would have prevented the whole thing, and it is by a wide margin the cheapest item in this account. Not for the runtime — a few minutes is plenty — but because it takes the mains transient and the servers never see it.
Beyond that: two members lost in one instant is not the same as two lost over months, and it is worth saying so before assuming the worst. Where one event takes several disks together, those disks are frequently recoverable.
What would have made this unrecoverable is precisely what the controller invites you to do — insert replacements and rebuild. Power the array down, leave the disks in their bays, label them if you must move them, and get the data secured before anything is written. RAID and server recovery is from £500 +VAT after a free 48-hour diagnostic.
Bring the drive to our Oxford Street reception, or post it over — it costs nothing to learn what went wrong. You’ll have a written figure from the fixed bands before any work begins.