Data Recovery Case File · NAS & Network Storage · Configured Is Not Verified
Redundancy That Was Set Up Is Not the Same as Redundancy That Held
This enquiry comes from someone who knows exactly what should have happened and can see that it did not. A large drive from a distributed filesystem, no longer visible to its host and suspected of controller failure, which "appears to have unique data chunks on it, despite the system being configured to keep more than one copy of every chunk." The configuration was correct and the outcome was not, which is the failure mode distributed storage is least good at announcing.
| Media | 24TB enterprise drive from a distributed filesystem cluster — not enumerating on its host; holding data chunks with no surviving replica elsewhere in the cluster |
| Reported situation | Enterprise drive forming part of a distributed filesystem · cluster configured to retain multiple copies of every data chunk · drive no longer visible as a device to the host operating system · controller failure suspected by the administrator · drive found to hold chunks without replicas elsewhere · recovery of the drive sought |
| Fault class | Device-level failure on a cluster member holding unreplicated chunks — replication incomplete despite configuration; recovery of this member the only route to the affected data |
| Equipment used | Replication gap accepted as making this member the sole source · device addressed through hardware with imposed timeouts rather than the host stack · controller state assessed in vendor technological modes · imaged write-blocked at the block level with chunk boundaries preserved · recovered chunks returned in a form the cluster can reingest |
The decode: how a configured replica fails to exist, and what follows
How distributed filesystems of this kind work: files are divided into chunks and each chunk is written to several members. A target replication count is set, and the system works towards it continuously rather than guaranteeing it at every instant.
Why that distinction is the whole of this case: the target is a goal, not an invariant. A chunk written recently, or written while the cluster was short of space or members, may exist in fewer copies than configured until the system catches up.
What can prevent it catching up: insufficient free space, members offline, a rebalancing operation not completing, or simply the volume of new data outpacing replication. Each leaves chunks temporarily under-replicated, and temporarily can be a long time.
Why nobody notices: the system reports its configuration, and everything reads normally as long as one copy exists. Under-replication is invisible during normal operation and only surfaces when a member is lost — which is exactly when it matters.
Why his observation is the valuable part: he checked which chunks were unique rather than assuming replication had worked. Most administrators discover this after deciding a failed member was expendable, and he has established it before doing so.
What that means for the drive: it is not a redundant member. For the affected chunks it is the sole copy, and it should be handled as a single-copy device rather than as one of many.
Why his suspicion of controller failure is plausible: the drive is not visible as a device at all. Absence from the operating system's device list places the fault below every filesystem and cluster question, and on an enterprise drive the electronics are the commonest cause of that.
Why that is comparatively good news: a controller that cannot present the drive does not alter what the platters hold. The chunks are intact behind a device that cannot introduce itself.
What the deliverable should be: a block-level image with chunk boundaries preserved, so the recovered content can be returned to the cluster rather than merely extracted. The cluster can reingest chunks it recognises, which is far better than delivering files.
What is worth doing across the rest of the cluster afterwards: an audit of replication state rather than of configuration. The question is which chunks currently have fewer copies than intended, and it is answerable before the next member fails.
On the bench
The replication gap was accepted as making this member the sole source — distributed filesystems dividing files into chunks written to several members and working continuously towards a target replication count rather than guaranteeing it at every instant, so chunks written recently or during space or member shortages may exist in fewer copies until the system catches up, which is invisible while any copy remains readable. Imaging ran write-blocked at the block level with chunk boundaries preserved.
The outcome
The replication gap accepted as making this member the sole source, the device addressed through hardware, and imaged with chunk boundaries preserved for reingestion. Free assessment, one fixed written figure including VAT; where a chip has to be removed, 50% of parts and labour is payable upfront with the balance only on success — otherwise no recovery, no fee. The decode: your replication target is a goal the system works towards, not an invariant it guarantees. Under-replication is invisible while one copy reads — and surfaces exactly when a member is lost.
When a cluster member holds data that should have been replicated
Treat that drive as a single-copy device rather than a redundant member, and don't dispose of it on the assumption the cluster has another copy. Distributed filesystems work continuously towards a replication target rather than guaranteeing it at every instant, so chunks written recently, or during a space or member shortage, can exist in fewer copies until the system catches up — and that's invisible during normal operation because everything reads fine while one copy survives. Afterwards, audit actual replication state rather than configuration, before the next member fails.
Treat it as a single copy — call Manchester Data Recovery on 0161 871 0788; replication gap accepted as making it the sole source, addressed through hardware, imaged with chunk boundaries preserved for reingestion.
Request a quote online →
Our case files are drawn from genuine enquiries received by our laboratory over the past ten years, anonymised to protect client confidentiality. Each one describes the diagnostic and recovery procedure our engineers apply to that fault, using the equipment listed.