Call us — 0161 871 0788
Mon–Fri · 9am–5:30pm · No fix, no fee
Start a free diagnostic →
// case file · SSD · Samsung EVO M.2 NVMe · thermal failure

Crashing when warm, then not at all.

A school’s external NVMe crashing after long transfers for weeks, then vanishing entirely. The flash was sound. The controller had been slowly cooked inside a sealed case — and the drive had been logging it the whole time.

← All case files · £300 + VAT, flat

Device

Samsung NVMe · external enclosure

Failure

Thermal — controller

Complication

Regulated data on minors

Outcome

Full recovery, 6 days

// the brief

What arrived, and what was at stake.

A school used an external Samsung NVMe drive for administrative storage — GCSE student records and exam data, the school database, class schedules, reports and spreadsheets.

It had been crashing intermittently for some weeks, always after a period of sustained use, and had latterly stopped being recognised altogether. The material included personal data on minors, which changes the handling requirements before any technical question is reached.

// what the drive was already telling them

NVMe records its own overheating.

This is the part worth knowing, because the evidence was available months before the failure and nobody had a reason to look at it.

Every NVMe drive maintains a SMART/health log the host can read at any time. Alongside the obvious figures — percentage used, data units written, unsafe shutdowns — it carries a composite temperature reading and, critically, two counters recording how long the drive has spent in thermal management state one and state two. Those are the throttling states: the drive slowing itself down deliberately to shed heat.

A drive with hours accumulated in those counters is a drive that has been running at its thermal limit routinely rather than occasionally. It is not a prediction of failure, but it is a direct measurement of the condition that produces this one, and it costs nothing to read.

The critical-warning byte in the same log has a dedicated bit for temperature threshold exceeded. On this drive it had been setting and clearing repeatedly — which is exactly what an intermittent crash after sustained use looks like from the drive’s side.

// on the bench

What the diagnosis found.

01

The failure pattern described heat, not wear

Crashes after sustained use, normal operation again once cool, then eventual total failure. Flash exhaustion does not behave that way — it degrades monotonically and shows in the percentage-used and available-spare figures, both of which were unremarkable here.

02

Physical evidence agreed

Discolouration of the board around the controller package, consistent with prolonged operation at temperature rather than one thermal event. Controllers in this class dissipate several watts from a package a few millimetres square.

03

The NAND tested sound

Which is the usual pattern. Heat fatigues the controller and its solder joints long before it damages the flash the controller manages — repeated expansion and contraction works on the ball-grid connections underneath the package, and an intermittent joint produces precisely this sequence.

04

The enclosure was half the fault

A compact external case with no meaningful thermal path, housing a drive designed to sit against a motherboard and shed heat into it.

External NVMe enclosures are where this fault concentrates, and the reason is structural rather than bad luck. An M.2 drive inside a laptop or desktop has airflow, a board to conduct into, and frequently a heatsink. Sealed in a small case with a USB bridge chip generating its own heat a centimetre away, it has none of that — and a sustained transfer will hold it above its throttling threshold for as long as the transfer runs.

// the recovery

How it was done.

The drive was taken out of its enclosure and addressed directly rather than over USB, which removed the bridge chip from the equation and allowed the controller to be reached in vendor-specific diagnostic mode.

That access is what the job turns on. An SSD does not store files at fixed physical addresses: the flash translation layer maps logical blocks to physical pages and moves them constantly for wear levelling and garbage collection. Those mapping tables live in a reserved service area, and without them the NAND is undifferentiated data — you can read every page and reconstruct nothing. Reading the service area first, while the controller was cool enough to respond, secured the map before anything else was attempted.

Imaging then ran in short passes with active cooling and rest intervals, because the fault that caused the failure also governs how the drive can be read: it was stable cold and unreliable warm, so each session ran until temperature climbed and then stopped. Regions that returned inconsistent data across passes were re-read until they settled rather than accepted first time — an intermittent solder joint will return plausible-looking rubbish as readily as it returns nothing.

Nothing was written to the drive at any stage.

// handling, because of what was on it

Personal data on pupils.

The technical work was the smaller half of this job. Because the material included personal information about children, the engagement ran under a signed data processing agreement naming the school as controller and us as processor, with a single named contact on each side.

All data remained in the UK throughout, working copies were held only for the duration of the job and destroyed on an agreed schedule with confirmation in writing, and access was limited to the engineer working the case. Where a recovery involves regulated records this is not an optional extra — the school has obligations that do not pause because a drive failed, and a recovery lab that cannot evidence its side of that is a liability rather than a supplier.

// outcome

What came back.

Everything. The controller had failed while the NAND remained sound, which is the usual outcome of thermal damage and the reason these recover as well as they do.

The school database was checked for structural consistency rather than merely copied — a database file that is byte-complete can still refuse to attach if it was captured mid-transaction — and records were confirmed readable in the software that needed them before the set was returned on fresh media.

// the transferable bit

What to take from this.

An external NVMe enclosure is a thermal compromise, and a cheap one is a bad thermal compromise. If you use one for sustained work, buy a case with an actual heatsink and a thermal pad contacting the controller, and check the drive’s thermal management counters occasionally — any free NVMe utility will show them.

Crashes that occur only after a period of use and clear up once the device has cooled are the warning sign, and they are very easy to blame on the computer instead. And if a portable drive holds the only copy of anything, it should not.

SSD and NVMe recovery is from £300 +VAT after a free 48-hour diagnostic, and where the data includes personal or regulated records we sign a data processing agreement as standard.

// read next

Related.

// your turn

Lost something that matters? Free diagnosis, a fixed price, and no fix, no fee.

Bring the drive to our Oxford Street reception, or post it over — it costs nothing to learn what went wrong. You’ll have a written figure from the fixed bands before any work begins.