The only agent that thinks for itself

Autonomous Monitoring with self-learning AI built-in, operating independently across your entire stack.

Unlimited Metrics & Logs
Machine learning & MCP
5% CPU, 150MB RAM
3GB disk, >1 year retention
800+ integrations, zero config
Dashboards, alerts out of the box
> Discover Netdata Agents

Centralized metrics streaming and storage

Aggregate metrics from multiple agents into centralized Parent nodes for unified monitoring across your infrastructure.

Stream from unlimited agents
Long-term data retention
High availability clustering
Data replication & backup
Scalable architecture
Enterprise-grade security
> Learn about Parents

Fully managed cloud platform

Access your monitoring data from anywhere with our SaaS platform. No infrastructure to manage, automatic updates, and global availability.

Zero infrastructure management
99.9% uptime SLA
Global data centers
Automatic updates & patches
Enterprise SSO & RBAC
SOC2 & ISO certified
> Explore Netdata Cloud

Deploy Netdata Cloud in your infrastructure

Run the full Netdata Cloud platform on-premises for complete data sovereignty and compliance with your security policies.

Complete data sovereignty
Air-gapped deployment
Custom compliance controls
Private network integration
Dedicated support team
Kubernetes & Docker support
> Learn about Cloud On-Premises

Powerful, intuitive monitoring interface

Modern, responsive UI built for real-time troubleshooting with customizable dashboards and advanced visualization capabilities.

Real-time chart updates
Customizable dashboards
Dark & light themes
Advanced filtering & search
Responsive on all devices
Collaboration features
> Explore Netdata UI

Monitor on the go

Native iOS and Android apps bring full monitoring capabilities to your mobile device with real-time alerts and notifications.

iOS & Android apps
Push notifications
Touch-optimized interface
Offline data access
Biometric authentication
Widget support
> Download apps

The future of infrastructure observability

See our strategic direction across AI-native observability, full-stack signals, operational intelligence, and enterprise platform maturity.

AI-native observability
Full-stack signal coverage
Operational intelligence
Enterprise platform maturity
Agent releases every 6 weeks
Cloud continuous delivery
> Explore Product Roadmap

Best energy efficiency

True real-time per-second

100% automated zero config

Centralized observability

Multi-year retention

High availability built-in

Zero maintenance

Always up-to-date

Enterprise security

Complete data control

Air-gap ready

Compliance certified

Millisecond responsiveness

Infinite zoom & pan

Works on any device

Native performance

Instant alerts

Monitor anywhere

AI-native observability

Continuous delivery

Open source foundation

80% Faster Incident Resolution

AI-powered troubleshooting from detection, to root cause and blast radius identification, to reporting.

True Real-Time and Simple, even at Scale

Linearly and infinitely scalable full-stack observability, that can be deployed even mid-crisis.

90% Cost Reduction, Full Fidelity

Instead of centralizing the data, Netdata distributes the code, eliminating pipelines and complexity.

See and Map Your Entire Network

Live topology, flow analytics, and SNMP device and trap monitoring — unified with your full-stack observability.

Control Without Surrender

SOC 2 Type 2 certified with every metric kept on your infrastructure.

Integrations

800+ collectors and notification channels, auto-discovered and ready out of the box.

800+ data collectors
Auto-discovery & zero config
Cloud, infra, app protocols
Notifications out of the box
> Explore integrations
Real Results
46% Cost Reduction

Reduced monitoring costs by 46% while cutting staff overhead by 67%.

— Leonardo Antunez, Codyas

Zero Pipeline

No data shipping. No central storage costs. Query at the edge.

From Our Users
"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

No Query Language

Point-and-click troubleshooting. No PromQL, no LogQL, no learning curve.

Enterprise Ready
67% Less Staff, 46% Cost Cut

Enterprise efficiency without enterprise complexity—real ROI from day one.

— Leonardo Antunez, Codyas

SOC 2 Type 2 Certified

Zero data egress. Only metadata reaches the cloud. Your metrics stay on your infrastructure.

Full Coverage
800+ Collectors

Auto-discovered and configured. No manual setup required.

Any Notification Channel

Slack, PagerDuty, Teams, email, webhooks—all built-in.

Built for the People Who Get Paged

Because 3am alerts deserve instant answers, not hour-long hunts.

Every Industry Has Rules. We Master Them.

See how healthcare, finance, and government teams cut monitoring costs 90% while staying audit-ready.

Monitor Any Technology. Configure Nothing.

Install the agent. It already knows your stack.
From Our Users
"A Rare Unicorn"

Netdata gives more than you invest in it. A rare unicorn that obeys the Pareto rule.

— Eduard Porquet Mateu, TMB Barcelona

99% Downtime Reduction

Reduced website downtime by 99% and cloud bill by 30% using Netdata alerts.

— Falkland Islands Government

Real Savings
30% Cloud Cost Reduction

Optimized resource allocation based on Netdata alerts cut cloud spending by 30%.

— Falkland Islands Government

46% Cost Cut

Reduced monitoring staff by 67% while cutting operational costs by 46%.

— Codyas

Real Coverage
"Plugin for Everything"

Netdata has agent capacity or a plugin for everything, including Windows and Kubernetes.

— Eduard Porquet Mateu, TMB Barcelona

"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

Real Speed
Troubleshooting in 30 Seconds

From 2-3 minutes to 30 seconds—instant visibility into any node issue.

— Matthew Artist, Nodecraft

20% Downtime Reduction

20% less downtime and 40% budget optimization from out-of-the-box monitoring.

— Simon Beginn, LANCOM Systems

Pay per Node. Unlimited Everything Else.

One price per node. Unlimited metrics, logs, users, and retention. No per-GB surprises.

Free tier—forever
No metric limits or caps
Retention you control
Cancel anytime
> See pricing plans

What's Your Monitoring Really Costing You?

Most teams overpay by 40-60%. Let's find out why.

Expose hidden metric charges
Calculate tool consolidation
Customers report 30-67% savings
Results in under 60 seconds
> See what you're really paying

Your Infrastructure Is Unique. Let's Talk.

Because monitoring 10 nodes is different from monitoring 10,000.

On-prem & air-gapped deployment
Volume pricing & agreements
Architecture review for your scale
Compliance & security support
> Start a conversation

Monitoring That Sells Itself

Deploy in minutes. Impress clients in hours. Earn recurring revenue for years.

30-second live demos close deals
Zero config = zero support burden
Competitive margins & deal protection
Response in 48 hours
> Apply to partner

Per-Second Metrics at Homelab Prices

Same engine, same dashboards, same ML. Just priced for tinkerers.

Community: Free forever · 5 nodes · non-commercial
Homelab: $90/yr · unlimited nodes · fair usage
> Get the Homelab Plan

$1,000 Per Referral. Unlimited Referrals.

Your colleagues get 10% off. You get 10% commission. Everyone wins.

10% of subscriptions, up to $1,000 each
Track earnings inside Netdata Cloud
PayPal/Venmo payouts in 3-4 weeks
No caps, no complexity
> Get your referral link
Cost Proof
40% Budget Optimization

"Netdata's significant positive impact" — LANCOM Systems

Calculate Your Savings

Compare vs Datadog, Grafana, Dynatrace

Savings Proof
46% Cost Reduction

"Cut costs by 46%, staff by 67%" — Codyas

30% Cloud Bill Savings

"Reduced cloud bill by 30%" — Falkland Islands Gov

Enterprise Proof
"Better Than Combined Alternatives"

"Better observability with Netdata than combining other tools." — TMB Barcelona

Real Engineers, <24h Response

DPA, SLAs, on-prem, volume pricing

Why Partners Win
Demo Live Infrastructure

One command, 30 seconds, real data—no sandbox needed

Zero Tickets, High Margins

Auto-config + per-node pricing = predictable profit

Homelab Ready
Free Video Course

8-episode Netdata tutorial by LearnLinux.tv

76k+ GitHub Stars

3rd most starred monitoring project

Worth Recommending
Product That Delivers

Customers report 40-67% cost cuts, 99% downtime reduction

Zero Risk to Your Rep

Free tier lets them try before they buy

AI Support Assistant, Available 24/7

Nedi has access to all official documentation, source code, and resources. Ask any question about Netdata—responds in your language.

Deployment & configuration
Troubleshooting & sizing
Alerts & notifications
Evidence-based answers
> Ask Nedi now

Never Fight Fires Alone

Docs, community, and expert help—pick your path to resolution.

Learn.netdata.cloud docs
Discord, Forums, GitHub
Premium support available
> Get answers now

60 Seconds to First Dashboard

One command to install. Zero config. 850+ integrations documented.

Linux, Windows, K8s, Docker
Auto-discovers your stack
> Read our documentation

76,000+ Engineers Strong

615+ contributors. 1.5M daily downloads. One mission: simplify observability.

Per-Second. 90% Cheaper. Data Stays Home.

Side-by-side comparisons: costs, real-time granularity, and data sovereignty for every major tool.

See why teams switch from Datadog, Prometheus, Grafana, and more.

> Browse all comparisons
Edge-Native Observability, Born Open Source
Per-second visibility, ML on every metric, and data that never leaves your infrastructure.
Founded in 2016
615+ contributors worldwide
Remote-first, engineering-driven
Open source first
> Read our story
Promises We Publish—and Prove
12 principles backed by open code, independent validation, and measurable outcomes.
Open source, peer-reviewed
Zero config, instant value
Data sovereignty by design
Aligned pricing, no surprises
> See all 12 principles
Edge-Native, AI-Ready, 100% Open
76k+ stars. Full ML, AI, and automation—GPLv3+, not premium add-ons.
76,000+ GitHub stars
GPLv3+ licensed forever
ML on every metric, included
Zero vendor lock-in
> Explore our open source
Build Real-Time Observability for the World
Remote-first team shipping per-second monitoring with ML on every metric.
Remote-first, fully distributed
Open source (76k+ stars)
Challenging technical problems
Your code on millions of systems
> See open roles
Meet the Team Behind Netdata
Conferences, meetups, and tradeshows where you can see Netdata in action and talk to the engineers who build it.
Live demos and deep dives
Book 1-on-1 meetings
Talks and panel sessions
Event recaps and photos
> See all events
Talk to a Netdata Human in <24 Hours
Sales, partnerships, press, or professional services—real engineers, fast answers.
Discuss your observability needs
Pricing and volume discounts
Partnership opportunities
Media and press inquiries
> Book a conversation
Your Data. Your Rules.
On-prem data, cloud control plane, transparent terms.
Trust & Scale
76,000+ GitHub Stars

One of the most popular open-source monitoring projects

SOC 2 Type 2 Certified

Enterprise-grade security and compliance

Data Sovereignty

Your metrics stay on your infrastructure

Validated
University of Amsterdam

"Most energy-efficient monitoring solution" — ICSOC 2023, peer-reviewed

ADASTEC (Autonomous Driving)

"Doesn't miss alerts—mission-critical trust for safety software"

Community Stats
615+ Contributors

Global community improving monitoring for everyone

1.5M+ Downloads/Day

Trusted by teams worldwide

GPLv3+ Licensed

Free forever, fully open source agent

Why Join?
Remote-First

Work from anywhere, async-friendly culture

Impact at Scale

Your work helps millions of systems

Troubleshooting

Linux Load Average Spikes: IO Wait & Bottlenecks

Learn to interpret load average correctly and use tools like iostat and vmstat to find the root cause of system performance issues
by Netdata Team · August 19, 2025

An alert fires: “High load average on production server.” Your heart rate quickens. You SSH into the machine and run a command like top, only to be confused. The CPU usage is hovering at 10%, but the load average is sky-high. What’s going on? If the CPU isn’t busy, what is the system “loaded” with? This common scenario highlights one of the most misunderstood metrics in Linux performance troubleshooting: the system load average.

Contrary to popular belief, load average is not a direct measure of CPU utilization. It’s a measure of demand for CPU resources. A high load average can signal a CPU bottleneck, but it can also be a symptom of a system struggling with slow I/O, excessive context switching, or other resource contention. Understanding the difference is the key to quickly diagnosing performance problems instead of chasing red herrings. This guide will demystify the Linux load average, explore its common causes, and provide a practical workflow for pinpointing the real bottleneck.

Demystifying Linux Load Average Beyond CPU Usage

When you run commands like uptime or top, you see three numbers representing the load average over the last 1, 5, and 15 minutes. They provide a trendline: if the one-minute average is higher than the fifteen-minute average, the load is increasing. If it’s lower, the load is decreasing. But what exactly is being averaged?

The Linux kernel calculates load average based on the number of processes that are either running or in an uninterruptible sleep state.

  • Runnable Processes (R state): These are processes that are actively using the CPU or are ready and waiting in the run queue for their turn on the CPU.
  • Uninterruptible Sleep Processes (D state): These are processes that are blocked, waiting for a resource to become available, most commonly disk or network I/O. They cannot be interrupted, even by a signal, until the I/O operation completes.

This is the crucial distinction. A system with a load average of 20 could have 20 processes all trying to max out the CPU, or it could have 20 processes all stuck waiting for a slow NFS mount to respond. In the second case, CPU utilization could be near zero, yet the system is heavily loaded.

CPU Utilization vs. Load Average

Think of it like a grocery store checkout.

  • CPU Utilization is how busy the cashier is. If they are scanning items 90% of the time, utilization is 90%.
  • Load Average is the number of people in the checkout line plus the person currently being served.

If you have one cashier (a single-core CPU) and a load average of 1.00, the cashier is perfectly busy. If the load average is 2.00, there’s one person being served and one person waiting. If the load is 0.50, the cashier is idle half the time. On a 4-core system, a load average of 4.00 means all cores are fully utilized. A load average of 8.00 on that same 4-core system means there’s a significant queue of tasks waiting for CPU time.

However, if the people in line are all waiting for a price check (our I/O wait analogy), the cashier isn’t busy, but the line is still long. This is how you get high load with low CPU utilization.

The Prime Suspects Investigating the Causes of High Load

When your load average spikes, the cause typically falls into one of three categories. Your job is to determine which one you’re dealing with.

CPU-Bound Bottlenecks Too Much Work, Not Enough CPU

This is the most straightforward cause. You simply have more active processes than your CPUs can handle. This leads to resource contention as processes compete for CPU time slices.

Symptoms:

  • The load average is high and correlates with the number of CPU cores (e.g., a load of 10 on an 8-core machine).
  • In command-line tools, the CPU usage is high, particularly the user space (%us) and kernel space (%sy) values. The I/O wait (%wa) is low.

Tools for Diagnosis:

  • Tools like top or htop, sorted by CPU, immediately show which processes are consuming the most CPU cycles.
  • The pidstat command provides a rolling update of CPU usage per process, which can be more useful than top’s snapshots.

If you identify a single process, like a database or application server, hogging the CPU, the next step is to use application-specific profiling tools to understand what it’s doing.

I/O Wait The Silent Killer of Performance

This is the most common cause of the confusing “high load, low CPU” problem. The system isn’t slow because the CPUs are overworked; it’s slow because processes are stuck waiting for slow hardware. This could be a failing hard drive, a saturated network link, or an overloaded storage array.

Symptoms:

  • The load average is high, but CPU usage (%us + %sy) is low.
  • In tools like top or iostat, the %wa or %iowait value is high. This is the percentage of time the CPU was idle but had at least one pending I/O request.

Tools for Diagnosis:

  • The vmstat command gives a great overview. The wa column under the cpu section shows I/O wait. The b column under the procs section shows the number of processes in uninterruptible sleep (the D state). A high number of blocked processes and a high percentage of I/O wait is a clear I/O bottleneck.
  • The iostat command is essential for drilling into disk performance. It provides stats per block device. Look for device utilization (%util). If this is near 100%, the disk is saturated. Also check await, the average time (in milliseconds) for I/O requests to be served. High values indicate the disk is struggling to keep up.
  • Once iostat confirms a disk bottleneck, iotop can show you which processes are generating the most disk read/write activity, much like top does for CPU.

The Overhead of Excessive Context Switching

A context switch occurs when the kernel switches the CPU from one process or thread to another. While this is a normal part of a multitasking operating system, an excessively high rate of context switching is pure overhead. The CPU spends its time saving and loading process states instead of doing useful work.

Symptoms:

  • Load average may be high, and CPU usage is dominated by system time (%sy). This is because the scheduler, part of the kernel, is working overtime.

Tools for Diagnosis:

  • The cs column in vmstat shows the number of context switches per second. There’s no single “bad” number; you need to establish a baseline for your system. A sudden, dramatic increase from the baseline is a red flag. The in column shows interrupts per second, which can be a cause of high context switching.
  • The pidstat tool is the best tool for finding the source. It shows context switches per process: voluntary context switches (cswch/s) and involuntary context switches (nvcswch/s). A process with a very high number of involuntary switches is often a sign of CPU pressure, while high voluntary switches might point to I/O issues.

A Practical Troubleshooting Workflow

When an alert for high system load average hits, stay calm and follow a logical process.

  1. Assess the Load: Is the load truly high for this system? A load of 4.0 on a 2-core machine is a problem. A load of 4.0 on a 32-core machine is trivial.
  2. Characterize the Problem: Run vmstat for 5-10 seconds. This is your command center. Look at the CPU columns first. Is user or system time high? You likely have a CPU-bound problem. Is I/O wait high? You have an I/O bottleneck. Are both low, but the context switch column is unusually high compared to its baseline? You may have a context switching issue. Check the processes columns. A high number in the runnable column confirms CPU pressure. A high number in the blocked column confirms an I/O problem.
  3. Drill Down with the Right Tool:
    • If CPU-bound: Use top or pidstat to find the process or processes consuming the most CPU.
    • If I/O-bound: Use iostat to identify the specific disk that is saturated. Then, use iotop to see which process is hammering that disk.
    • If Context Switching: Use pidstat to identify the process with the highest rate of context switches.

By following this workflow, you can move from a vague “high load” alert to a specific, actionable root cause in minutes.

The Linux load average is a powerful but nuanced metric. By understanding that it represents demand from both running and blocked processes, you can avoid the common trap of only looking at CPU utilization. High load is a symptom, not a diagnosis. Learning to use tools like vmstat, iostat, and pidstat allows you to look past the symptom and uncover the real bottleneck, whether it’s an exhausted CPU, a struggling disk, or a system drowning in its own overhead.

Monitoring these metrics over time is key to identifying deviations from the norm. Netdata automatically collects thousands of system metrics, including load average, I/O wait, and context switches, displaying them on real-time, interactive dashboards. This allows you to spot trends and anomalies instantly, without needing to manually run commands. Get started with Netdata for free and gain immediate visibility into the performance of your entire infrastructure.