Call us — 0161 871 0788
Mon–Fri · 9am–5:30pm · No fix, no fee
Start a free diagnostic →
// case file · Server · HP ProLiant · RAID 1 · cache battery

Two error messages. One underlying cause.

An HP ProLiant refusing to boot, reporting a logical drive error and placing the array into interim recovery mode — alongside a cache battery fault the client had been ignoring for months as a separate nuisance. Those two are the same story far more often than they look.

← All case files · from £500 + VAT

Device

HP ProLiant · Smart Array RAID 1

Failure

Lost writes · logical drive error

Complication

Cache battery failed months earlier

Outcome

Full recovery, 6 days

// the brief

What arrived, and what was at stake.

The server had been reporting a cache battery fault for some time and running normally otherwise, so it went on the list and stayed there. Then it failed to boot, reporting a logical drive error and announcing that the array had entered interim recovery mode.

The client attempted to halt that to regain access, which did not work, and the server stayed down with the whole business on it. Neither disk had failed.

// why the battery matters

A cache battery warning is a data warning.

01

Writes are acknowledged before they reach the platters

Smart Array controllers carry a write-back cache. The operating system is told a write has completed as soon as it lands in that cache, not when it reaches the disks. That is precisely why these controllers are quick, and it is the whole point of the feature.

02

The battery is what protects the writes in flight

If power is lost, the battery (or on later units a supercapacitor backing flash) preserves the cache contents so the controller can flush them to the platters when the machine comes back. Without it, anything sitting in cache at the moment of interruption simply evaporates.

03

A failed battery is supposed to disable write-back

The controller normally falls back to write-through mode, where nothing is acknowledged until it is genuinely on disk. That is slower, and it is the sluggishness people notice and complain about — but it is the controller protecting them.

04

Lost writes leave the file system describing itself incorrectly

Any interruption while write-back is still active loses data the operating system firmly believes was committed. Metadata updates go missing while the file data lands, or the reverse. That is one of the commonest routes to a logical drive error on a server with no failed disk at all.

The order of events here is the ordinary one. Battery fails. Somebody schedules the replacement for the next maintenance window. Between those two points there is a window in which acknowledged writes have nothing standing behind them, and a power event lands in it.

// interim recovery mode

A status, not a process.

This caused most of the lost time here, so it is worth being exact. Interim recovery mode is HP’s term for a logical drive running with reduced or absent redundancy — the controller declaring that it can no longer guarantee normal protection. On a two-disk RAID 1 it means the volume is effectively running on one member.

It is a description of a state, not a repair in progress. There is nothing to halt and nothing to wait out, which is what the client spent time attempting. Left running, the array is doing ordinary work with no margin whatsoever: any read error on the surviving member is now a data loss rather than something parity or the mirror quietly absorbs.

The useful response to seeing it is to stop and secure the data, in that order.

// the recovery

How it was done.

Both drives were imaged read-only before anything else. Neither had failed mechanically — SMART was unremarkable on both, with no reallocation activity and no pending sectors, entirely consistent with a logical fault from lost writes rather than a dying disk.

The mirror was then reassembled from the images, entirely outside the ProLiant. The controller and its metadata were treated as unreliable evidence throughout, on the straightforward grounds that the controller was implicated in the failure: a cache that lost writes may also have lost metadata updates describing those writes, so its account of the array’s state was taken as a hypothesis rather than as fact and checked against the disks.

Resolving disagreement between the members. This is the substance of the job. On a mirror that lost writes, the two disks do not agree everywhere — a write reached one member and not the other before the interruption. Choosing arbitrarily, or simply trusting whichever the controller considers current, produces a volume that mounts and is quietly wrong.

Every divergent region was resolved against the file system’s own structures instead. Where NTFS metadata pointed at a version that was internally consistent — the record complete, the run list resolving to allocated clusters, the allocation bitmap agreeing — that version was taken. Where the journal held a record of the transaction, it decided the question. The result is a volume whose contents are consistent with its own metadata rather than merely present.

The file system was then repaired from the assembled volume, and the SQL databases on it checked for structural consistency page by page rather than assumed good because they attached.

// outcome

What came back.

Everything, in six days, because the lost writes were recent and limited in extent. Had the server run in write-back mode with a dead battery for months rather than weeks, the answer would have been materially different — every power event in that period contributes another set of divergences, and at some point the file system stops being repairable into anything trustworthy.

// the transferable bit

What to take from this.

A battery warning on a caching controller is not a maintenance note for next quarter. It means the protection around in-flight writes has gone, and the failure it produces is a file system that no longer describes itself correctly — on hardware where every disk is in perfect health.

Two practical points follow. If you see the warning and cannot replace the part immediately, disable write-back caching in the controller configuration; the server will be noticeably slower and it will stop putting data at risk. And if a logical drive error appears on a server whose disks are all healthy, resist the instinct to let the controller try to sort it out — the controller is frequently the reason you are there.

Server and RAID recovery is from £500 +VAT after a free 48-hour diagnostic.

// read next

Related.

// your turn

Lost something that matters? Free diagnosis, a fixed price, and no fix, no fee.

Bring the drive to our Oxford Street reception, or post it over — it costs nothing to learn what went wrong. You’ll have a written price from our fixed bands before any work starts.