Technician comparing a solid state drive and a hard drive beside an open desktop tower on a repair bench

Diagnostic guides

Diagnostic Guides for Failing and Used Drives

These walkthroughs are for the moments when the numbers matter. A machine that has started to feel wrong, a used SSD with no warranty, a drive running hotter than it should. Each guide tells you what to read and what to do with the reading, in the order you would actually do it.

What this collection covers

Nine walkthroughs, grouped so you can jump straight to the situation you are in. Each one ends with the reading that matters most and how to act on it.

Guide one

How to tell a drive is dying

Start with attributes 05 and 197. A reallocated sectors count that is stable and small is usually fine; the same count growing across readings is not. Pending sectors are the sharper signal because they are unresolved. Then check self-test results and listen for the drive. A healthy drive should not click rhythmically under light load.

When in doubt, the order is always the same: copy data, then diagnose. Nothing in this guide is worth risking a file you cannot replace.

01

Read attribute 05 twice

Note the raw count today, then again in a week or two of normal use. Compare the trend, not the snapshot number.

02

Watch pending sectors

Attribute 197 counts sectors the drive could not read and has not resolved. A rising number is hard to argue with.

03

Run a short self-test

The drive's own short test catches read failures on the surface before the counter does. Run it while the machine is otherwise idle.

04

Listen under light load

Rhythmic clicking that repeats with every read is a mechanical warning. A healthy drive stays quiet when the workload is easy.

Copy the data first. Diagnose second. The drive can wait; the files may not.

Jump to reading the data
Hands holding a used solid state drive up to window light beside a laptop and a stack of drives on a desk mat

Guide two

How to check a used SSD before buying

Ask for a screenshot of the drive's SMART data, not a photo of the label. Look at wear level or life-left percentage, total bytes written, and power-on hours, in that order. A drive with high write volume and low remaining life is a poor bet regardless of price.

Confirm the vendor so the attribute IDs match what you are reading. An SSD with clean wear but unstable error counters is also a reason to walk away.

Wear and life left

Percent of rated endurance already spent. Sort out which value the vendor's tool prints before comparing drives.

Total bytes written

How much data has actually moved through the drive. Compare this against the model's endurance rating, not against a year count.

Power-on hours

Age, not health. Useful as context once you know how the drive was used, and required if you want the calculator to score it.

Guide three

Why 60°C is a problem

Heat accelerates wear on both HDD platters and SSD NAND. A drive that idles hot, not just under load, is usually fighting airflow or a failing fan. 60°C sustained is a problem; the fix is often a case change or a fan, not a new drive. Check whether temperature is reported as attribute 194 or 190 so you read the right field.

Power-on hours: what the number actually tells you

Hours measure age, not health. A NAS drive in always-on service may log tens of thousands of hours and remain within spec. What you want to know is whether the drive is being asked to do more than it was built for. Vendor line and workload rating tell you that; the raw hour count alone does not.

Read the temperature you actually selected

Which field belongs to which attribute depends on the drive family. Confirm it before you compare.

Attribute 194
Temperature in Celsius on the drive that reports a single current reading.
Attribute 190
Airflow temperature. Reported on some drives as a separate field from the drive's own reading.
Attribute 231
Some SSD families log remaining life here instead. Read the vendor's own notes before you trust a label.
See the attribute reference
Person examining a bare hard drive at a kitchen table with a laptop pushed aside and notes face down

Guide four

Reading smartctl output without guessing

smartctl prints raw and normalized values side by side. The raw value is the number that grows; the normalized value is the drive's own score. Read the raw values for error counters and the normalized values for the pass/fail comparison against the threshold. This one habit avoids most misreadings.

Raw

What the drive has actually counted — sectors reassigned, hours run, bytes moved. This is the number to track over time.

Normalized

The drive's own score, set so higher usually means healthier. This is what most tools compare against the threshold.

Threshold

The line the drive declares for itself. Below it, the value is treated as failed, which is why the raw reading alone is not the verdict.

Worst

The lowest normalized value the drive has ever reported. A Worst score near the threshold flags drift even when the current score still looks fine.

Common beginner mistakes

None of these require special tools to avoid. They just require you to not trust a single reading as a verdict.

  • Treating one high reading as the verdict

    A single spike can come from a bad cable, a warm case or one bad write. Look at whether the number is growing.

  • Ignoring trend in favor of a snapshot

    Two readings a week apart tell you more about a dying hard drive than any single smartctl report.

  • Trusting a clean SMART report as a guarantee

    A clean report is a clean report for today. It tells you nothing about the load you are about to put on the drive.

  • Replacing the drive before backing up

    Moving failing media from one bay to another is not a backup. Copy first, then swap the hardware.

Bare hard drive on a workbench beside a coiled SATA cable and screwdriver under a work lamp

Guide five

Uncorrectable and CRC errors

Offline uncorrectable sectors (198) and UDMA CRC errors (199) point in different directions. Uncorrectable errors are the drive failing to read its own data. CRC errors are usually a cabling or controller problem. Clear the cable first if CRC is the only signal growing.

Attribute 198

Offline uncorrectable sectors

The drive could not read a block during its own offline scan. It is a data-integrity signal, not a connection signal. If this counter is moving, plan the copy-out.

Most useful when it moves between two readings of the same drive.

Attribute 199

UDMA CRC error count

The cable or controller introduced a transfer error. The drive usually treats it as a retry. Replace the SATA cable and reseat the power connector before you blame the drive.

If the counter stops moving after the cable swap, the drive is likely still doing its job.

Questions we get before people start

Short answers to the ones that keep coming up. Where a guide above covers the same ground, the full reading is there.

Which reading tells me the most in one look?

Pending sectors. They are unresolved read failures the drive has not corrected yet, so they are a sharper signal than a reallocated count that stays flat.

Do I need the same tool for every drive brand?

Not for reading the counters — the attribute IDs and meanings are stable across the major families. You do want the vendor's own utility for firmware-level details and for tools that label attributes by name instead of ID.

My drive is running at 60°C. Is it dead?

No, but it is worth fixing. Sustained heat at that level accelerates wear. Check airflow, dust and fan speed first; a case change or a new fan is often enough to bring it back down.

How do I use the reading with the calculator?

Enter the drive type, the counters you have, and the temperature and hours. The calculator scores what you give it — it does not replace a backup. If you are unsure of a field, the attribute reference explains what each one means.

Quiet dusk home office with an open desktop tower and a spare solid state drive resting on an anti-static mat

Next step: score the drive you are worried about

Open the calculator, enter the counters and readings you already have, and compare the result against the thresholds in the attribute reference. If a number here still looks wrong to you, send a correction and it will be reviewed.

Spotted an error in a guide or an attribute description? Send a correction and the team at 452 Silicon Drive, Suite 102, San Jose, CA 95110 will review it. Weekday support is available on +1-408-555-0192.