Call us — 0161 871 0788
Mon–Fri · 9am–5:30pm · No fix, no fee
Start a free diagnostic →
← All case files // case file · NAS / RAID

Two amber lights, and a layout nobody documents.

A photographer’s Drobo showing two amber bays at once. BeyondRAID is proprietary and undocumented, so no standard tool reads it — and that, rather than the two failed disks, is the actual difficulty.

DeviceDrobo · BeyondRAID
FaultTwo members lost together
Turnaround11 days
Outcome96% recovered
MethodMap reconstruction · imaging

What arrived

The unit held a working photography archive: camera raw files, edited exports and delivered client galleries going back years. Two bays lit amber within a short period of each other and the volume stopped mounting. The owner had been through Drobo Dashboard, which declared the array unable to protect data and offered no route forward. They stopped there and rang. Stopping was the single most useful thing they did.

Why a Drobo is a harder problem than a failed RAID 5

Conventional RAID has published behaviour, and that is what makes it tractable. A RAID 5 stripes data with parity rotating through one of a small number of documented patterns — left-symmetric, left-asymmetric and their right-hand equivalents — at one of a handful of common chunk sizes. A recovery tool can therefore generate candidate geometries, apply each to the images, and test the result: valid file system structures appear at plausible offsets, or they do not. The search space is small and the answer is verifiable.

BeyondRAID is not a stripe layout at all. Drobo virtualises storage below the file system: incoming data is broken into regions and distributed across whatever disks happen to be present, with redundancy applied per region rather than uniformly across an array, and a mapping layer recording where every logical block physically lives. That is what allows mixed disk sizes, hot expansion and single-to-dual redundancy switching on a live volume. It is genuinely good engineering.

It is also entirely undocumented. There is no published specification, no standard tool that reads it, and — critically — no way to derive the layout by testing conventional stripe patterns, because it does not use any. Guessing geometry is not a slow route here; it is not a route. The recovery depends on locating and parsing Drobo’s own mapping structures out of the disks themselves.

What had actually failed

Every member was imaged before anything else, including both flagged disks. Neither was mechanically destroyed, which turned out to matter more here than it would on a conventional array.

The first flagged disk carried a heavy population of unreadable sectors in tight bands rather than scattered across the surface. Banding at regular LBA intervals is the signature of damage confined to one or two heads: on a multi-platter drive, logical addresses are interleaved across surfaces, so a single bad head produces defects at a repeating stride rather than in one contiguous run. Mapping the failures against the drive’s head count confirmed it — two surfaces were affected, the rest read normally.

The second had dropped out under sustained load with a failing head that still read, slowly, when the drive was cool and the queue was shallow. It had not been ejected because it was dead; it had been ejected because it exceeded the enclosure’s response timeout under load.

On a RAID 5 that distinction would be academic once parity covers the gap. Under per-region redundancy it is decisive: some regions had a redundant copy and some did not, so every additional sector recovered from a failed member fed directly into the final result instead of being made redundant by parity held elsewhere. The imaging effort on those two disks was the recovery, not preparation for it.

Imaging strategy

Both damaged members were imaged in graduated passes rather than one long attempt. A first pass ran sequentially at a large block size with a short timeout, taking everything that returned without retry. A reverse pass followed, which frequently picks up sectors near a defect that a forward pass abandoned mid-retry. Only then were smaller block sizes applied to the remaining gaps, in short sessions with the drives rested and cooled between them, since head-related read failure worsens measurably as a drive warms.

On the banded disk, the heads serving the two damaged surfaces were disabled for the bulk passes so the drive could read the healthy surfaces at full speed without stalling on retries it was never going to satisfy, and re-enabled only for targeted attempts at the end. The Drobo enclosure itself was never powered again once the disks were out.

Reconstructing the map

With images in hand the work became structural. Drobo’s mapping metadata is written to the disks in a consistent form even though that form is not published, and it is identifiable: descriptor structures repeat at regular intervals, carry recognisable field patterns, and cross-reference between members in ways that ordinary file data does not. Located, the descriptors were parsed to establish which physical extents on which disks backed which logical block addresses, and how redundancy had been applied region by region.

That mapping was rebuilt in software and used to assemble the logical volume from the images alone. From the assembled volume the file system was rebuilt, and only then did anything resembling a folder tree exist.

Verification, which for a photographic archive means rendering

A raw file that opens is not the same as a raw file that is intact. Camera raw containers hold an embedded JPEG preview alongside the sensor data, and a file can present valid headers and a readable preview while the raw payload behind it is damaged — which is exactly the failure mode a partial recovery produces.

Every image was therefore decoded rather than opened: the raw payload demosaiced to confirm it processes end to end, and the result compared against the embedded preview to catch files where the two disagree. Files that failed that check were listed by name rather than quietly returned. The client received a manifest of what came back and what did not, before paying anything.

Outcome

About 96%. The shortfall was regions that were unreadable on both failed members at once, in areas where the per-region redundancy had already been consumed — the only part of the volume with no surviving copy anywhere on any disk.

Worth stating plainly, because Drobo owners ask: the proprietary layout made this job long and made it expensive. It did not make it impossible. What would have made it impossible was continuing to power the unit and letting it attempt its own relayout across two failing members.

// ready when you are

Facing something similar? Let's help.

Kick off with an instant online quote, or ring us and talk it through first. Either way you’ll know a clear, fixed price before any work starts.

Peter House, Oxford Street, Manchester M1 5AN · Mon–Fri 9am–5:30pm · No fix, no fee on most jobs