Call us — 0161 871 0788
Mon–Fri · 9am–5:30pm · No fix, no fee
Start a free diagnostic →

Data Recovery Case File · NAS & Network Storage · The Rebuild Was the Risk

Rebuilding Reads Every Surviving Member From End to End

His enquiry contains the action that most often turns a recoverable array into a lost one. A four-bay unit in a parity configuration where "one drive has completely failed, one drive intermittently fails", and a fresh drive was fitted to rebuild — without success. A rebuild is the heaviest possible demand on the remaining members, and it was asked of a set that already contained one drive known to be unreliable.

MediaFour-bay network storage unit in a single-parity configuration — one member failed outright, one member failing intermittently; rebuild attempted with a replacement drive
Reported situationFour-bay unit configured with single parity across four members · one member completely failed · one member failing intermittently · replacement drive fitted · rebuild attempted · rebuild unsuccessful · two replacement drives available · recovery of the array sought
Fault classRebuild attempted with a second member unreliable — full-surface read demanded of a degraded set; array state after the attempt to be assessed before any further action
Equipment usedRebuild attempt assessed for what it wrote before any further action · each member imaged individually write-blocked · intermittent member imaged under capped timeouts with regions revisited across passes · array parameters derived from the members rather than the unit · array assembled offline from images with no rebuild permitted on originals

The decode: what a rebuild demands, and what the attempt may have done

What rebuilding actually involves: reconstructing the missing member's entire contents. Every block of the replacement drive is calculated from the corresponding blocks on all the surviving members, which means reading every one of them from beginning to end without a gap.

Why that is the heaviest operation an array ever performs: normal use reads scattered fragments on demand. A rebuild is a continuous full-surface read of every remaining drive, for hours, and it cannot skip anything because every block is needed.

Why a drive that fails intermittently is the worst possible participant: the rebuild will reach its bad regions, because it reaches everything. Ordinary use may never have touched them, which is why the drive seemed merely intermittent rather than failed.

What happens when it fails mid-rebuild: the array loses a second member while already reconstructing the first. With single parity there is no information left to continue from, and the rebuild stops — which is what he observed.

Why the attempt may have cost more than time: a rebuild writes to the replacement drive throughout, and depending on the unit may update metadata on the surviving members. What was written and how far it got needs establishing before anything else is done.

Why the intermittent drive is now the critical component: its recoverable extent decides the outcome. Parity fills in the failed member wherever the others read, so the only regions that cannot be reconstructed are those where the intermittent drive also fails.

Why hours of rebuild reading may have worsened exactly that: the drive spent that period working continuously through regions it struggles with. An intermittent drive subjected to a full-surface read is a drive being pushed through its worst areas repeatedly, which is how intermittent becomes permanent.

What should have happened instead, and it is worth stating for others: the array should have been imaged before any rebuild. A rebuild is safe when every remaining member is healthy and reckless when one is not — and the unit will offer it either way, because it does not assess the survivors first.

What is done now: every member imaged individually under capped timeouts, the intermittent one revisited across multiple passes, and the array assembled offline from the images with parameters derived from the members themselves.

What must not happen: no further rebuild attempts, no fitting the second replacement drive, and no initialising the array. The unit will keep offering to rebuild, and each attempt spends the intermittent drive further.

On the bench

The rebuild attempt was assessed for what it wrote before any further action — reconstruction calculating every block of a replacement from corresponding blocks on all surviving members, requiring a continuous full-surface read of each with nothing skipped, which reaches regions ordinary use never touches and therefore surfaces an intermittent member's worst areas. Each member was imaged individually, the intermittent one under capped timeouts with regions revisited across passes, and the array assembled offline from images.

The outcome

The rebuild attempt assessed first, each member imaged individually, and the array assembled offline with no rebuild permitted on originals. Free assessment, one fixed written figure including VAT, charged per drive, with 50% of parts and labour upfront and the balance only on successful recovery. The decode: a rebuild reads every surviving member end to end, which is why it found your intermittent drive's bad regions — ordinary use never reached them. That drive's recoverable extent now decides the outcome.

Before rebuilding an array with a member you already doubt

Image the array first, and don't fit the second replacement. A rebuild reconstructs every block of the new drive from the corresponding blocks on all the survivors, so it reads each of them end to end for hours with nothing skipped — which is how it reaches the bad regions ordinary use never touched. That makes an intermittently failing member the worst possible participant: the rebuild will find its weak areas, and hours of continuous reading is how intermittent becomes permanent. The unit will offer to rebuild regardless, because it doesn't assess the survivors first.

Rebuild failed with a second drive already unreliable?
Don't try again — call Manchester Data Recovery on 0161 871 0788; rebuild attempt assessed first, each member imaged individually under capped timeouts, array assembled offline from images.
Request a quote online →

Our case files are drawn from genuine enquiries received by our laboratory over the past ten years, anonymised to protect client confidentiality. Each one describes the diagnostic and recovery procedure our engineers apply to that fault, using the equipment listed.