Call us — 0161 871 0788
Mon–Fri · 9am–5:30pm · No fix, no fee
Start a free diagnostic →

Data Recovery Case File · NAS & Network Storage · Still Working Is the Window

An Array Still Serving Is the Best Position You Will Have From Here

His enquiry describes a system that is working and should not be assumed to keep working. A four-drive production server in a mirrored-and-striped configuration where "one drive has failed and another is failing, and the system at present is working and essential to our business." Still working is the whole of the opportunity — and the request he makes, to have the drives replaced, is not the first thing that should happen.

MediaFour-drive array in a mirrored-and-striped configuration on a production server — one member failed and a second indicating a fault; array currently serving
Reported situationProduction server in business-critical use · four-drive array in a mirrored and striped configuration · one member indicated as failed · a second member indicating a fault · system currently operating and serving · owner requesting replacement of the failed drives
Fault classDegraded array with a second member indicating fault — tolerance dependent on which members are affected; rebuild carrying the principal risk while service continues
Equipment usedPair membership of the affected drives established before any replacement was considered · a complete backup taken from the running array as the first action · every member imaged individually write-blocked before any rebuild · array parameters derived from the members rather than the controller · replacement and rebuild deferred until content was secured

The decode: what this configuration tolerates, and what to do first

What a mirrored-and-striped set actually is: pairs of drives holding identical copies, with data spread across the pairs. Each pair is a mirror, and the pairs together give capacity and speed.

What that means for tolerance, and it is the crucial detail: the set survives losing one drive from each pair. It does not survive losing both drives of the same pair — so which two are affected matters far more than how many.

Why that must be established immediately: if the failed drive and the failing drive are in different pairs, the array is comfortable. If they are in the same pair, it is one event from total loss, and the two situations look identical from outside.

What the first action should be, and it is not a replacement: a complete backup taken from the array while it is still serving. A working array is readable at full speed through its own controller, which is faster, cheaper and more complete than any recovery.

Why that ordering matters more than anything else here: the array is currently in the best condition it will be in. Every hour of production use is more reading on a member already indicating a fault, and the opportunity narrows rather than widens.

Why replacing a drive is the risky step rather than the safe one: fitting a replacement triggers a rebuild. A rebuild reads the surviving member of that pair from end to end without pause — and if that member is the one indicating a fault, the rebuild is the operation most likely to finish it.

Why that risk is highest in exactly this situation: a drive that has been indicating a fault has regions it struggles with. Ordinary use may never have touched them; a full-surface read will.

What the right sequence therefore is: establish pair membership, back up completely while the system serves, image any member indicating a fault, and only then replace and rebuild. The rebuild becomes a routine operation once nothing depends on it succeeding.

Why we are not the right people to do the replacement itself: that is server maintenance rather than recovery, and it belongs with whoever supports the hardware. What is useful from us is the assessment, the imaging and the sequence — and saying so plainly is more helpful than taking the work.

What should not happen while any of this is arranged: no drives pulled to inspect them, and no rebuild started. A running degraded array should be left running and read, not interfered with.

On the bench

Pair membership of the affected drives was established before any replacement was considered — a mirrored-and-striped set comprising mirrored pairs with data distributed across them, surviving the loss of one member per pair but not both members of the same pair, so which drives are affected governs tolerance rather than how many. A complete backup was taken from the running array as the first action, a serving array being readable at full speed through its own controller.

The outcome

Pair membership established first, a complete backup taken while the array served, and every member imaged before any rebuild. Free assessment, one fixed written figure including VAT, charged per drive where recovery is required. The decode: the system still working is the opportunity, not the reassurance. Back it up now while it serves — and establish whether the two affected drives are in the same pair, because that decides how close this is.

A production array with a second drive going

Back it up completely before replacing anything — a serving array reads at full speed through its own controller, which is faster, cheaper and more complete than any recovery, and it is currently in the best condition it will be in. Establish which pair each affected drive belongs to, since a mirrored-and-striped set survives losing one member per pair but not both of the same pair, and that matters more than the count. Understand why replacement is the risky step: fitting a drive triggers a rebuild, which reads the surviving member of that pair end to end — the operation most likely to finish a drive already indicating a fault.

Live array with a second drive indicating a fault?
Back it up before replacing anything — call Manchester Data Recovery on 0161 871 0788; pair membership established first, backup taken while the array serves, members imaged before any rebuild.
Request a quote online →

Our case files are drawn from genuine enquiries received by our laboratory over the past ten years, anonymised to protect client confidentiality. Each one describes the diagnostic and recovery procedure our engineers apply to that fault, using the equipment listed.