Almost every array that reaches us lost its first disk weeks before anyone noticed — because redundancy keeps everything working while quietly spending the margin. By the time the volume disappears, the failure people remember is the second one. Here is what each level actually survives, and why a rebuild is the wrong response.
$ mdr diagnose /dev/raid → Array: RAID 5 · 6 × 4 TB · 20 TB volume → Status: OFFLINE — 2 disks failed, rebuild failed → Client: confidential · Manchester M2 $ mdr engineer-working → Member disks: all 6 imaged read-only → Parameters: order + stripe + parity solved → Array: rebuilt virtually from images $ mdr verify → ✓ databases — 412 GB → ✓ shares + VMs — 17.8 TB → ✓ array recovered — data back
A rebuild writes to every surviving member for hours, and where it cannot read parity cleanly it writes wrong data. It is the single commonest way an array that would have recovered fully becomes fragments. Power it down and leave the disks in their bays.
The level tells you exactly how much has to go wrong before the array stops — and how it behaves on the way there.
Almost every array that arrives here lost a disk long before anyone was told.
That is the awkward consequence of redundancy: the array keeps serving data with a member down, so nothing looks wrong. The alert goes to a mailbox nobody reads, or the amber light is on a unit in a cupboard. Weeks pass. Then a second disk goes and the volume disappears in an afternoon, and the failure everyone remembers is the second one.
Which is why the degraded state matters more than the failure. An array running without redundancy has no margin left, and every hour of continued use is an hour at full risk. If a controller reports a disk down, the useful response is to copy the data off before replacing anything — not to order a disk and carry on.
The controller is the one component that has already demonstrated it cannot be relied on.
Including any disk already marked failed — an ejected member frequently still holds most of its data and is often the difference between partial and full recovery.
Stripe size, disk order, parity rotation and offset established by testing candidates until reassembly produces coherent file-system structures rather than noise.
The array is rebuilt from the images, entirely outside the hardware. Nothing is written to any original disk at any stage.
Where one member has an unreadable sector, the remainder computes it. Only stripes with two unreadable members in the same place cannot be resolved.
NTFS, ReFS, ext4, XFS or Btrfs reconstructed from the assembled volume, and databases checked for structural consistency rather than merely copied.
The practical part, and the part that most affects how long it takes.
Send every disk including the failed one, each labelled with its bay number — 1, 2, 3 and so on. Bay order is central to reconstruction and while it can be derived by analysis, having it costs you nothing and saves real time. For a hardware controller, send the card as well if it comes out easily.
Then tell us three things: the controller or NAS model, the level if you know it, and — most importantly — what has already been attempted. A rebuild that was started and abandoned, an array re-created to “test” it, or disks reinserted in a different order all change the approach completely, and we would far rather know than discover it. Nobody is judged for it; we simply need the accurate picture.
A two-disk mirror and a twelve-disk array are not the same job, and the quote reflects that.
RAID, NAS and server recovery starts from £500 +VAT, confirmed after the free 48-hour diagnostic. Enterprise SAN environments are priced separately from £1,250 +VAT after a scoping call. Where members need physical drive-level work there is a 50% deposit, with the balance payable only on a successful recovery.
Business terms are standard: NDAs signed as a matter of course, one named contact from diagnostic to delivery, invoicing rather than card payment, and data remaining in the UK throughout.
Get the array to us, and tell us the RAID level, disk count and controller. The diagnostic that follows costs nothing, and one figure goes in writing before any work begins.
Nothing starts until the array to us — all its disks, labelled in order, or the whole server is on the bench. Pack it properly and put your contact details in with it. The diagnostic that follows costs nothing, and one figure goes in writing before any work begins.
Posting it? Use something tracked and insured — whatever is on the drive is worth considerably more than the postage. Bringing it in? Weekdays, 9am to 5:30pm, and it still wants packing as above for the journey.
Not sure it is worth sending? Tell us about the array, RAID level and fault, plus anything already attempted — an engineer reads every one of these and comes back with an honest view of the odds and a price band, before you post anything.
We’ll be in touch shortly. If it’s urgent, call 0161 871 0788.
The level tells you the arithmetic. The platform tells you where the description of that arithmetic lives — and that is what usually goes missing.
Hardware controllers — Dell PERC, HP Smart Array, LSI and Broadcom, Adaptec, IBM ServeRAID. These hold the array configuration on the card, sometimes with a copy on the disks and sometimes not. A firmware update or a card failure can leave healthy disks and no description of how they fit together, which is a recoverable situation but not one the card can fix itself.
Software arrays — Linux mdadm, Windows dynamic disks and Storage Spaces, ZFS. The configuration lives on the members, which helps, though superblocks updated mid-rebuild can describe a state that no longer exists.
Appliances — Synology, QNAP, Buffalo, Drobo and the rest add their own layers above the array: SHR, thin-provisioned pools, BeyondRAID. All are reconstructable, but the layers have to be unwound in order rather than treated as one volume.
The questions we’re asked most about recovering a RAID array.
Often, yes. What matters is whether the second disk failed completely or simply dropped out — an ejected member is usually still largely readable. Both come in, every member is imaged, and the array is reconstructed from the copies rather than the hardware.
Not if the data matters. A rebuild reads every sector of every survivor for hours and writes to the members as it goes; where parity cannot be read cleanly it writes wrong data. It is the commonest way a recoverable array becomes an unrecoverable one.
Yes. Degraded means the redundancy is already spent, so any further fault means total loss. Copy the data off before replacing anything. The comfortable feeling of it still working is exactly what makes this state expensive.
Yes, including any already marked failed. Every level except a mirror needs the full set to reconstruct the volume, and even an ejected disk usually contributes readable data. Label them by bay before removing them.
That is a normal starting point rather than a problem. Stripe size, member order, parity rotation and offset are all derived by analysing the disks themselves — the controller’s own metadata is treated as a hint, not a fact.
From £500 plus VAT, fixed in writing after a free 48-hour diagnostic. Enterprise SAN work is priced separately from £1,250 plus VAT. Physical drive-level work carries a 50% deposit with the balance due only on success.
Every one of these writes its own metadata format, and none of them is trusted during a recovery — the geometry is derived from the data on the disks and checked against parity, because the controller is usually implicated in the failure.
From £500 +VAT, quoted once we have seen the members. Every disk imaged read-only and the array rebuilt away from the controller that lost it — and if the unit is still powered and offering to rebuild, switch it off before anything else.