Data Recovery Case File · NAS & Network Storage · Two Counts, One Answer Needed
On Double-Parity, Whether Two or Three Have Failed Is the Entire Question
Her enquiry identifies a discrepancy that decides everything. A twelve-drive array in a double-parity configuration "with either two or three disks failed — the hardware lights disagree with the firmware." Double parity tolerates two failures and not three, so the array is either degraded and recoverable in place or beyond its own ability to reconstruct — and nothing else about the case matters until that is settled.
| Media | Twelve-drive array in double-parity configuration within a network unit — failed member count reported inconsistently between indicators and controller |
| Reported situation | Twelve-drive array configured with double parity · either two or three members reported failed · hardware indicators and controller firmware reporting different counts · owner seeking assessment of what is possible · business context with limited on-site availability |
| Fault class | Multiple member failure at or beyond parity tolerance — reconstruction feasibility governed by the true count and by partial readability of failed members |
| Equipment used | True member state established by direct assessment rather than by controller or indicator report · every member imaged individually write-blocked · marginal members imaged under capped timeouts with regions revisited across passes · array parameters derived from the members rather than the controller · array assembled offline with no rebuild permitted on originals |
The decode: why the counts differ, and why neither should be trusted
What double parity provides: the ability to lose any two members and continue. Information is distributed so that the contents of any two drives can be reconstructed from the remaining ten — genuine protection, and considerably stronger than single parity.
What it does not provide: tolerance of a third. At three failures the mathematics runs out, and the array cannot reconstruct itself from what remains.
Why the two reports differ: they measure different things. Indicators generally reflect what the enclosure's own monitoring concluded, while the controller reports which members it has ejected from the array — and a drive can be marked failed by one and not the other.
Why a controller ejects drives conservatively: it removes a member on a timeout or an error threshold rather than on a diagnosis. A drive that responded slowly once may be marked failed while remaining substantially readable, which is very common and very important here.
Why that is the hopeful reading of the discrepancy: the third drive may not be failed in any meaningful sense. Ejected is a decision the controller made, not a statement about the hardware — and a member that reads well is a member that counts.
Why neither count should be relied on: both are the unit's own view. The actual condition of twelve drives is established by assessing twelve drives, and it frequently differs from what any controller believes.
Why that assessment must come before anything else: the natural next step is to replace a drive and let the array rebuild. On an array at or beyond its tolerance, a rebuild reads every surviving member end to end and can take further members with it.
What is done instead: every member imaged individually, marginal ones under capped timeouts with difficult regions revisited across passes, and the array assembled offline from those images with parameters derived from the members themselves.
Why partial readability of a failed member is often enough: parity fills in what is missing wherever the others read. A third drive readable across most of its surface leaves only the regions where it and the others coincide unreconstructible — which is frequently a small proportion.
Why the array must not be brought back into service meanwhile: it is a live system in a business, and the pressure to restore it is exactly what causes the loss. Nothing should be replaced, rebuilt or initialised until the members have been read.
On the bench
True member state was established by direct assessment rather than by controller or indicator report — double parity tolerating two failures and not three, while indicators reflect enclosure monitoring and the controller reports ejected members, a controller ejecting on timeout or error threshold rather than on diagnosis, so a member marked failed may remain substantially readable. Every member was imaged individually, marginal ones under capped timeouts, and the array assembled offline with no rebuild permitted on originals.
The outcome
Member state established by direct assessment, every member imaged individually, and the array assembled offline from images. Free assessment, one fixed written figure including VAT, charged per drive, with 50% of parts and labour upfront and the balance only on successful recovery. The decode: the disagreement may be good news. A controller ejects members on timeouts and error thresholds rather than on diagnosis — so the third drive may be readable, and a readable member counts.
An array where the reported failure count is unclear
Don't replace anything or start a rebuild until every member has been read. On double parity the difference between two and three failures is the difference between degraded and beyond self-reconstruction, so that count decides everything — and neither the indicators nor the controller is authoritative, since they measure different things. A controller ejects members on timeouts and error thresholds rather than diagnosis, so a drive marked failed may be substantially readable, and a readable member still counts. A rebuild at or beyond tolerance reads every survivor end to end.
Don't rebuild it — call Manchester Data Recovery on 0161 871 0788; member state established by direct assessment, every member imaged individually, array assembled offline from images.
Request a quote online →
Our case files are drawn from genuine enquiries received by our laboratory over the past ten years, anonymised to protect client confidentiality. Each one describes the diagnostic and recovery procedure our engineers apply to that fault, using the equipment listed.