The only agent that thinks for itself

Autonomous Monitoring with self-learning AI built-in, operating independently across your entire stack.

Unlimited Metrics & Logs
Machine learning & MCP
5% CPU, 150MB RAM
3GB disk, >1 year retention
800+ integrations, zero config
Dashboards, alerts out of the box
> Discover Netdata Agents

Centralized metrics streaming and storage

Aggregate metrics from multiple agents into centralized Parent nodes for unified monitoring across your infrastructure.

Stream from unlimited agents
Long-term data retention
High availability clustering
Data replication & backup
Scalable architecture
Enterprise-grade security
> Learn about Parents

Fully managed cloud platform

Access your monitoring data from anywhere with our SaaS platform. No infrastructure to manage, automatic updates, and global availability.

Zero infrastructure management
99.9% uptime SLA
Global data centers
Automatic updates & patches
Enterprise SSO & RBAC
SOC2 & ISO certified
> Explore Netdata Cloud

Deploy Netdata Cloud in your infrastructure

Run the full Netdata Cloud platform on-premises for complete data sovereignty and compliance with your security policies.

Complete data sovereignty
Air-gapped deployment
Custom compliance controls
Private network integration
Dedicated support team
Kubernetes & Docker support
> Learn about Cloud On-Premises

Powerful, intuitive monitoring interface

Modern, responsive UI built for real-time troubleshooting with customizable dashboards and advanced visualization capabilities.

Real-time chart updates
Customizable dashboards
Dark & light themes
Advanced filtering & search
Responsive on all devices
Collaboration features
> Explore Netdata UI

Monitor on the go

Native iOS and Android apps bring full monitoring capabilities to your mobile device with real-time alerts and notifications.

iOS & Android apps
Push notifications
Touch-optimized interface
Offline data access
Biometric authentication
Widget support
> Download apps

The future of infrastructure observability

See our strategic direction across AI-native observability, full-stack signals, operational intelligence, and enterprise platform maturity.

AI-native observability
Full-stack signal coverage
Operational intelligence
Enterprise platform maturity
Agent releases every 6 weeks
Cloud continuous delivery
> Explore Product Roadmap

Best energy efficiency

True real-time per-second

100% automated zero config

Centralized observability

Multi-year retention

High availability built-in

Zero maintenance

Always up-to-date

Enterprise security

Complete data control

Air-gap ready

Compliance certified

Millisecond responsiveness

Infinite zoom & pan

Works on any device

Native performance

Instant alerts

Monitor anywhere

AI-native observability

Continuous delivery

Open source foundation

80% Faster Incident Resolution

AI-powered troubleshooting from detection, to root cause and blast radius identification, to reporting.

True Real-Time and Simple, even at Scale

Linearly and infinitely scalable full-stack observability, that can be deployed even mid-crisis.

90% Cost Reduction, Full Fidelity

Instead of centralizing the data, Netdata distributes the code, eliminating pipelines and complexity.

See and Map Your Entire Network

Live topology, flow analytics, and SNMP device and trap monitoring — unified with your full-stack observability.

Control Without Surrender

SOC 2 Type 2 certified with every metric kept on your infrastructure.

Integrations

800+ collectors and notification channels, auto-discovered and ready out of the box.

800+ data collectors
Auto-discovery & zero config
Cloud, infra, app protocols
Notifications out of the box
> Explore integrations
Real Results
46% Cost Reduction

Reduced monitoring costs by 46% while cutting staff overhead by 67%.

— Leonardo Antunez, Codyas

Zero Pipeline

No data shipping. No central storage costs. Query at the edge.

From Our Users
"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

No Query Language

Point-and-click troubleshooting. No PromQL, no LogQL, no learning curve.

Enterprise Ready
67% Less Staff, 46% Cost Cut

Enterprise efficiency without enterprise complexity—real ROI from day one.

— Leonardo Antunez, Codyas

SOC 2 Type 2 Certified

Zero data egress. Only metadata reaches the cloud. Your metrics stay on your infrastructure.

Full Coverage
800+ Collectors

Auto-discovered and configured. No manual setup required.

Any Notification Channel

Slack, PagerDuty, Teams, email, webhooks—all built-in.

Built for the People Who Get Paged

Because 3am alerts deserve instant answers, not hour-long hunts.

Every Industry Has Rules. We Master Them.

See how healthcare, finance, and government teams cut monitoring costs 90% while staying audit-ready.

Monitor Any Technology. Configure Nothing.

Install the agent. It already knows your stack.
From Our Users
"A Rare Unicorn"

Netdata gives more than you invest in it. A rare unicorn that obeys the Pareto rule.

— Eduard Porquet Mateu, TMB Barcelona

99% Downtime Reduction

Reduced website downtime by 99% and cloud bill by 30% using Netdata alerts.

— Falkland Islands Government

Real Savings
30% Cloud Cost Reduction

Optimized resource allocation based on Netdata alerts cut cloud spending by 30%.

— Falkland Islands Government

46% Cost Cut

Reduced monitoring staff by 67% while cutting operational costs by 46%.

— Codyas

Real Coverage
"Plugin for Everything"

Netdata has agent capacity or a plugin for everything, including Windows and Kubernetes.

— Eduard Porquet Mateu, TMB Barcelona

"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

Real Speed
Troubleshooting in 30 Seconds

From 2-3 minutes to 30 seconds—instant visibility into any node issue.

— Matthew Artist, Nodecraft

20% Downtime Reduction

20% less downtime and 40% budget optimization from out-of-the-box monitoring.

— Simon Beginn, LANCOM Systems

Pay per Node. Unlimited Everything Else.

One price per node. Unlimited metrics, logs, users, and retention. No per-GB surprises.

Free tier—forever
No metric limits or caps
Retention you control
Cancel anytime
> See pricing plans

What's Your Monitoring Really Costing You?

Most teams overpay by 40-60%. Let's find out why.

Expose hidden metric charges
Calculate tool consolidation
Customers report 30-67% savings
Results in under 60 seconds
> See what you're really paying

Your Infrastructure Is Unique. Let's Talk.

Because monitoring 10 nodes is different from monitoring 10,000.

On-prem & air-gapped deployment
Volume pricing & agreements
Architecture review for your scale
Compliance & security support
> Start a conversation

Monitoring That Sells Itself

Deploy in minutes. Impress clients in hours. Earn recurring revenue for years.

30-second live demos close deals
Zero config = zero support burden
Competitive margins & deal protection
Response in 48 hours
> Apply to partner

Per-Second Metrics at Homelab Prices

Same engine, same dashboards, same ML. Just priced for tinkerers.

Community: Free forever · 5 nodes · non-commercial
Homelab: $90/yr · unlimited nodes · fair usage
> Get the Homelab Plan

$1,000 Per Referral. Unlimited Referrals.

Your colleagues get 10% off. You get 10% commission. Everyone wins.

10% of subscriptions, up to $1,000 each
Track earnings inside Netdata Cloud
PayPal/Venmo payouts in 3-4 weeks
No caps, no complexity
> Get your referral link
Cost Proof
40% Budget Optimization

"Netdata's significant positive impact" — LANCOM Systems

Calculate Your Savings

Compare vs Datadog, Grafana, Dynatrace

Savings Proof
46% Cost Reduction

"Cut costs by 46%, staff by 67%" — Codyas

30% Cloud Bill Savings

"Reduced cloud bill by 30%" — Falkland Islands Gov

Enterprise Proof
"Better Than Combined Alternatives"

"Better observability with Netdata than combining other tools." — TMB Barcelona

Real Engineers, <24h Response

DPA, SLAs, on-prem, volume pricing

Why Partners Win
Demo Live Infrastructure

One command, 30 seconds, real data—no sandbox needed

Zero Tickets, High Margins

Auto-config + per-node pricing = predictable profit

Homelab Ready
Free Video Course

8-episode Netdata tutorial by LearnLinux.tv

76k+ GitHub Stars

3rd most starred monitoring project

Worth Recommending
Product That Delivers

Customers report 40-67% cost cuts, 99% downtime reduction

Zero Risk to Your Rep

Free tier lets them try before they buy

AI Support Assistant, Available 24/7

Nedi has access to all official documentation, source code, and resources. Ask any question about Netdata—responds in your language.

Deployment & configuration
Troubleshooting & sizing
Alerts & notifications
Evidence-based answers
> Ask Nedi now

Never Fight Fires Alone

Docs, community, and expert help—pick your path to resolution.

Learn.netdata.cloud docs
Discord, Forums, GitHub
Premium support available
> Get answers now

60 Seconds to First Dashboard

One command to install. Zero config. 850+ integrations documented.

Linux, Windows, K8s, Docker
Auto-discovers your stack
> Read our documentation

76,000+ Engineers Strong

615+ contributors. 1.5M daily downloads. One mission: simplify observability.

Per-Second. 90% Cheaper. Data Stays Home.

Side-by-side comparisons: costs, real-time granularity, and data sovereignty for every major tool.

See why teams switch from Datadog, Prometheus, Grafana, and more.

> Browse all comparisons
Edge-Native Observability, Born Open Source
Per-second visibility, ML on every metric, and data that never leaves your infrastructure.
Founded in 2016
615+ contributors worldwide
Remote-first, engineering-driven
Open source first
> Read our story
Promises We Publish—and Prove
12 principles backed by open code, independent validation, and measurable outcomes.
Open source, peer-reviewed
Zero config, instant value
Data sovereignty by design
Aligned pricing, no surprises
> See all 12 principles
Edge-Native, AI-Ready, 100% Open
76k+ stars. Full ML, AI, and automation—GPLv3+, not premium add-ons.
76,000+ GitHub stars
GPLv3+ licensed forever
ML on every metric, included
Zero vendor lock-in
> Explore our open source
Build Real-Time Observability for the World
Remote-first team shipping per-second monitoring with ML on every metric.
Remote-first, fully distributed
Open source (76k+ stars)
Challenging technical problems
Your code on millions of systems
> See open roles
Meet the Team Behind Netdata
Conferences, meetups, and tradeshows where you can see Netdata in action and talk to the engineers who build it.
Live demos and deep dives
Book 1-on-1 meetings
Talks and panel sessions
Event recaps and photos
> See all events
Talk to a Netdata Human in <24 Hours
Sales, partnerships, press, or professional services—real engineers, fast answers.
Discuss your observability needs
Pricing and volume discounts
Partnership opportunities
Media and press inquiries
> Book a conversation
Your Data. Your Rules.
On-prem data, cloud control plane, transparent terms.
Trust & Scale
76,000+ GitHub Stars

One of the most popular open-source monitoring projects

SOC 2 Type 2 Certified

Enterprise-grade security and compliance

Data Sovereignty

Your metrics stay on your infrastructure

Validated
University of Amsterdam

"Most energy-efficient monitoring solution" — ICSOC 2023, peer-reviewed

ADASTEC (Autonomous Driving)

"Doesn't miss alerts—mission-critical trust for safety software"

Community Stats
615+ Contributors

Global community improving monitoring for everyone

1.5M+ Downloads/Day

Trusted by teams worldwide

GPLv3+ Licensed

Free forever, fully open source agent

Why Join?
Remote-First

Work from anywhere, async-friendly culture

Impact at Scale

Your work helps millions of systems

$ guides / smartctl
SMARTCTL · DISK HEALTH PLAYBOOK

The drive reports on its own health — but PASSED lies, raw values mislead, and some failures never show up at all

smartctl reads the firmware's own instrumentation: remapped sectors, NAND wear, interface errors, temperature, mechanical wear. We trace what each signal actually means, why the overall health check misses roughly a third of failures, and what to do when one specific attribute or drive state starts to move.

"

SMART gets you a health check in a single command, then hands you a set of numbers that are easy to read and just as easy to misread.

The one-liner is a trap. smartctl -H returns PASSED right up until a drive carrying hundreds of Current_Pending_Sector and a climbing Reallocated_Sector_Ct finally trips a generous vendor threshold — and Google's 2007 fleet study found 36% of failed drives showed no SMART warning at all. UDMA_CRC_Error_Count spikes and a team swaps a perfectly healthy drive when the real fault was the cable. Raw_Read_Error_Rate reads in the billions on a Seagate and looks catastrophic when it is normal by design. An NVMe drive quietly flips to read-only (Critical Warning bit 3) with Available Spare at zero. A controller dies and the device simply vanishes from the bus, leaving a clean SMART history behind it.

These guides are written for engineers who already run disks in production, not for people learning what a sector is. The goal is the mental model of what the firmware is actually measuring, which signals lead a failure and which only confirm it, why rate-of-change beats absolute thresholds every time, and the runbook for the specific attribute or state you just watched move.

How disk health actually reaches you

S.M.A.R.T. is not a monitoring agent — it is the drive's firmware reporting on itself, and smartctl only reads what the firmware already knows. Health flows up a stack from the physical media to the operating system, and most surprises live between the layers: a signal that leads failure on one layer, a summary verdict that lags it on another, and whole failure modes that bypass SMART entirely.

01
host + kernel I/O
What the operating system experiences: <code>iostat</code> await, <code>/proc/diskstats</code>, and kernel messages like <code>I/O error, dev sdX</code>. The host sees failures — controller hangs, command timeouts, SCSI error recovery — that never appear in SMART. smartctl reads the drive; the kernel reads the experience.
HOST
02
transport / interface
The physical path between host and drive: SATA/SAS cable, connector, backplane, HBA, PCIe lanes. Corruption here shows as <code>UDMA_CRC_Error_Count</code> (ATA) or SAS PHY counters, and can downshift the link to a slower speed. The media is fine; the wire is not.
LINK
03
drive controller + firmware
The drive's brain: ECC on every read, the SSD flash-translation layer, garbage collection, wear leveling — and SMART itself. When this dies the drive vanishes from the bus, and SMART can give no warning, because the component that reports it is the one that failed.
CONTROLLER
04
SMART attributes + health log
The firmware's self-report. ATA exposes vendor-specific attribute IDs with normalized VALUE, RAW, and THRESH columns; NVMe exposes a standardized health log; the binary PASSED/FAILED verdict summarizes both — optimistically, against thresholds tuned to limit warranty claims.
SMART
05
spare pool + remapping
The finite reserve that keeps a degrading drive alive. <code>Current_Pending_Sector</code> awaits a verdict, <code>Reallocated_Sector_Ct</code> records remaps already done, <code>Offline_Uncorrectable</code> counts the ones that failed; on SSD/NVMe it is <code>Available Spare</code>. Exhaust it and the next bad block is data loss.
SPARE
06
NAND / platter media
Where the bits physically live. SSD NAND has finite program/erase cycles tracked as <code>Percentage Used</code> and wear-leveling count; HDD magnetic surface degrades into bad sectors and rising read-error rates. This layer wears out predictably on an SSD and erratically on an HDD.
MEDIA
07
mechanical assembly (HDD)
Spindle motor, actuator arm, and heads. <code>Spin_Retry_Count</code>, <code>Spin_Up_Time</code>, <code>Seek_Error_Rate</code>, and <code>G-Sense_Error_Rate</code> report wear and shock. None of it exists on an SSD — and a single failure here can end the drive instantly rather than gradually.
MECHANICAL
08
environment
Temperature, power, and vibration act on every layer above, often exponentially. Heat degrades NAND retention and HDD bearings; unstable power drives unsafe shutdowns and spin retries across many drives at once; vibration causes seek errors. The context that decides how fast everything else fails.
ENVIRONMENT

Why this matters: 'the disk is slow' or 'I'm seeing disk errors' can be a bad cable retransmitting, pending sectors forcing read retries, an SSD stalling on garbage collection, an NVMe drive thermally throttling, a spindle struggling to spin up, or a controller about to disappear. The symptom rhymes but each layer has a different signal — and a different fix. And the overall health check can read PASSED through most of them.

The failures you'll actually see

Most disk incidents fall into a small set of recurring shapes. Recognise the shape and triage — replace the drive, or the cable, or the power supply — gets dramatically faster.

CRITICAL

The slow-dying drive

Bad sectors accumulate. Current_Pending_Sector appears, Reallocated_Sector_Ct climbs as the spare pool is consumed, and every read that lands on a suspect sector triggers a multi-second retry. The system looks frozen — high iostat await with an idle CPU — while the drive dies in slow motion over days.

  • Current_Pending_Sector (ID 197) non-zero and rising
  • Reallocated_Sector_Ct (ID 5) climbing week over week
  • iostat await spiking to seconds with the CPU idle (high iowait)
  • ATA error log showing UNC errors at specific LBAs
Investigate
CRITICAL

The SSD wear-out cliff

SSDs degrade gradually, then fail abruptly. Available Spare falls toward its threshold, Percentage Used passes 100%, and media errors begin — then the controller flips the drive to read-only, or dead, within hours. Unlike an HDD, the step from slightly worn to completely gone can be a single day.

  • Available Spare at or below the vendor threshold (typically 10%)
  • Percentage Used at or above 100%
  • Media and Data Integrity Errors starting to increment
  • Wear_Leveling_Count / Media_Wearout_Indicator near end of scale (SATA SSD)
Investigate
ACTIVE

The bad cable, not the drive

UDMA_CRC_Error_Count climbs and throughput drops, but every media attribute is zero. The corruption is on the wire — a loose or failing cable, backplane port, or HBA — and CRC catches it, so data integrity holds while the link may downshift to a slower speed. Replace the drive and the new one develops the same errors.

  • UDMA_CRC_Error_Count (ID 199) increasing, often just after maintenance
  • SATA link at current: below its maximum negotiated speed
  • Reallocated (5), Pending (197), and Uncorrectable (198) all zero
  • SAS: rising invalid-dword and loss-of-sync PHY counters
Investigate
CRITICAL

The drive that vanished

The device node was present and is now gone. A controller, PCB regulator, connector, or backplane failed — and because the component that reports SMART is the one that died, there was no warning. The SMART history is clean right up to the disappearance.

  • /dev/sdX or /dev/nvmeXnY present before, absent now
  • dmesg showing device removal, bus reset, or error-recovery failure
  • Clean SMART on the last successful collection
  • RAID array reporting a degraded or missing member
Investigate
IMMINENT

PASSED, but dying

smartctl -H returns PASSED while the drive carries hundreds of pending sectors, dozens of reallocations, and real latency. PASSED only means no pre-fail attribute has crossed a generous vendor threshold — and 36% of failed drives (Google 2007) showed no warning at all. Trusting the summary is how surprise failures happen.

  • Overall health PASSED with non-zero pending / reallocated / uncorrectable
  • An individual attribute VALUE approaching its THRESH
  • Counters rising that never yet trip the binary verdict
  • Kernel I/O errors on a drive SMART calls healthy
Investigate
CRITICAL

The spindle that won't spin up

Spin_Retry_Count goes non-zero — the motor failed to reach operating speed on the first attempt. A drive that struggles to spin up may not come back after the next power cycle. One drive means a failing motor; many drives at once means the power supply or the 12V rail, not the disks.

  • Spin_Retry_Count (ID 10) non-zero on a drive in service
  • Spin_Up_Time (ID 3) rising well above its own baseline
  • Multiple drives showing retries together (a power problem)
  • Unsafe-shutdown count rising alongside
Investigate
Choosing a tool

Best Disk Health & S.M.A.R.T. Monitoring Tools

A ranked review of the tools teams actually shortlist here, what each one is genuinely good at, and how the pricing behaves as you scale.

SMART monitoring maturity levels

Disk-health monitoring works in four practical levels. Each is a complete operation, not a stepping stone. Pick the level that matches how much the data on these drives matters. Anything holding real data should land at the second level.

Level 1: Survival

Know when a drive has died or declared itself dying

Survival monitoring is the floor: is the drive present, and does either its own firmware or its most basic attributes say it is failing? You will not predict failure, but you will not be blind to a drive that has already failed or called itself failed. Enough for a workstation or a disposable node; not enough for anything holding data you care about.

  • Device present Does /dev/sdX or /dev/nvmeXnY still exist? A vanished node is the drive gone.
  • SMART overall health (PASSED/FAILED) A FAILED verdict is the drive's own death notice — page on it.
  • Reallocated_Sector_Ct (ID 5) Any growth means the finite spare pool is being consumed.
  • Current_Pending_Sector (ID 197) Unreadable sectors awaiting remap; the active-degradation signal.
  • Drive temperature Sustained operation above the drive-type's rated maximum.
  • NVMe Critical Warning Any non-zero bit is a firmware-declared active problem.

Level 2: Operational

Catch drives on their way out, not just already out

Operational monitoring is what any team storing real data should run. Survival catches drives that already failed; operational catches the ones on their way. It watches the leading media and endurance signals, separates cable faults from drive faults, and confirms you can actually see every physical drive — the coverage that turns a 3am surprise into a planned replacement.

  • Offline_Uncorrectable (ID 198) Confirmed permanent data loss; alert on growth, key on the attribute name.
  • UDMA_CRC_Error_Count (ID 199) Rising CRC with zero media errors is a cable, not a drive.
  • NVMe Available Spare The remaining spare pool; below the vendor threshold is urgent.
  • NVMe Percentage Used Rated endurance consumed; watch the trend toward 100%.
  • NVMe Media and Data Integrity Errors Growth here is confirmed NAND-level corruption.
  • SMART blind-spot check Every expected physical drive actually returning SMART data.
  • Scheduled self-tests Short weekly, extended monthly — the drive will not scan itself.

Level 3: Mature

Read trends and logs before the summary ever changes

Mature monitoring is trend-based, not threshold-based. It tracks how fast each counter is moving, reads the error and self-test logs, projects SSD endurance runway from write volume, and watches the kernel's view alongside the drive's. This is where you catch the drive that is technically PASSED but visibly accelerating toward failure.

  • Rate-of-change on every counter 5 sectors gained this week beats 50 stable over five years.
  • ATA / NVMe error log New entries and their type (UNC, ICRC, CCTO), not just the count.
  • Self-test results Read/servo/electrical failures, and tests that never complete.
  • Write volume vs rated TBW Data Units Written / Total_LBAs_Written for endurance runway.
  • Power_On_Hours Fleet age; the denominator that normalizes every other attribute.
  • NVMe Composite Temperature Time Minutes spent overheating, even when the drive is cool now.
  • Host I/O errors (kernel logs) The failures the drive itself cannot see.
  • Unsafe shutdown count Rising counts flag power instability before it corrupts data.

Level 4: Expert

Reactive depth added after real incidents

Expert signals get added after a specific incident proved you needed them: a batch of drives that failed together, a firmware bug that bricked a model, an SSD burned out by a misconfigured workload. Most fleets never need all of these. Add the ones your own failure history points to — and always baseline on growth, not lifetime absolutes.

  • Seek_Error_Rate / Spin_Up_Time trends HDD mechanical wear, read via the normalized value only.
  • G-Sense_Error_Rate Shock and vibration actually reaching the drive.
  • Cross-drive correlation Same symptom on many drives at once = infrastructure, not disk.
  • Firmware + identity tracking Unexpected version changes, counter resets, known-buggy firmware.
  • First-observation baselining Alert on growth from a captured baseline, never lifetime absolutes.
  • Composite risk scoring Combine several declining attributes into one per-drive verdict.
  • SMART collection latency A drive slow to answer smartctl is an early controller-trouble sign.

Operating mistakes worth avoiding

The traps disk-monitoring teams keep falling into. Each has a clear, well-known fix. Most teams only learn it after a surprise failure or a needlessly replaced drive.

Trusting SMART PASSED as a clean bill of health

<code>smartctl -H</code> returns PASSED until a pre-fail attribute crosses a vendor threshold set to minimize warranty claims, not to protect your data. A drive can report PASSED with hundreds of pending sectors and seconds-long read latency, and Google's 2007 fleet study found 36% of failed drives triggered no SMART warning at all. Watch the individual attributes and the NVMe health log, never the summary alone.

Alerting on absolute values instead of rate of change

A <code>Reallocated_Sector_Ct</code> of 50 accumulated over five years is stable; a count that went from 0 to 5 this week is a drive actively dying. Most monitoring asks 'is the value above X?' and misses the acceleration that is the actionable signal. Track deltas per drive over time.

Blaming the drive for a cable problem

<code>UDMA_CRC_Error_Count</code> climbs, the team swaps the drive, and the replacement develops the same errors — because the cable, connector, or backplane port was always the fault. CRC errors with zero media errors (IDs 5/197/198) point at the transport layer. Check that relationship before condemning a disk.

Panicking at Seagate raw values

Seagate encodes <code>Raw_Read_Error_Rate</code> (ID 1) and <code>Seek_Error_Rate</code> (ID 7) as packed counters whose raw values routinely reach the billions and are completely normal. Alerting on raw ID 1 or ID 7 above zero produces constant false alarms — and operators who then learn to ignore SMART entirely. Only the normalized VALUE trending toward THRESH matters for these attributes.

Never scheduling extended self-tests

Drives do not scan themselves. Without a scheduled extended self-test, latent bad sectors in rarely-read data stay hidden until a scrub, backup, or restore finally reads them — turning a preventable warning into a production I/O error. A short weekly and an extended monthly test is the cheapest insurance there is.

Not tracking SSD write endurance

Teams deploy SSDs, never compare Data Units Written against the datasheet TBW, and are surprised when a drive dies early. It died exactly on schedule — the schedule was just never computed. Trend write volume and NVMe <code>Percentage Used</code> to know the replacement date months ahead.

The RAID controller and VM blind spot

Behind a hardware RAID controller, plain <code>smartctl</code> sees only the virtual device; inside a VM or cloud instance, SMART is usually absent or meaningless. Teams believe monitoring works while it sees nothing. Use <code>-d megaraid,N</code> or <code>-d cciss,N</code> passthrough on RAID, run SMART on the hypervisor host, and reconcile expected drives against monitored ones.

Paging on lifetime counters at first deployment

Roll SMART monitoring onto an existing fleet and every drive with any history — Offline_Uncorrectable, NVMe Media Errors, Unsafe Shutdowns — fires at once. Capture a baseline on the first scrape and alert only on growth from it. Absolute lifetime values at rollout are an alarm flood, not a signal.

SMART disk-health runbooks in this section

Each guide is a focused runbook for one attribute, state, or failure pattern. Pick one when an attribute just moved, or use the categories to learn the area.

WHERE TO GO NEXT

Setting up SMART monitoring, or staring at an attribute that just moved?

If you're starting from scratch, the monitoring checklist is the path of least regret. If a specific attribute or drive state just changed, jump straight to the runbook that matches it.