The only agent that thinks for itself

Autonomous Monitoring with self-learning AI built-in, operating independently across your entire stack.

Unlimited Metrics & Logs
Machine learning & MCP
5% CPU, 150MB RAM
3GB disk, >1 year retention
800+ integrations, zero config
Dashboards, alerts out of the box
> Discover Netdata Agents

Centralized metrics streaming and storage

Aggregate metrics from multiple agents into centralized Parent nodes for unified monitoring across your infrastructure.

Stream from unlimited agents
Long-term data retention
High availability clustering
Data replication & backup
Scalable architecture
Enterprise-grade security
> Learn about Parents

Fully managed cloud platform

Access your monitoring data from anywhere with our SaaS platform. No infrastructure to manage, automatic updates, and global availability.

Zero infrastructure management
99.9% uptime SLA
Global data centers
Automatic updates & patches
Enterprise SSO & RBAC
SOC2 & ISO certified
> Explore Netdata Cloud

Deploy Netdata Cloud in your infrastructure

Run the full Netdata Cloud platform on-premises for complete data sovereignty and compliance with your security policies.

Complete data sovereignty
Air-gapped deployment
Custom compliance controls
Private network integration
Dedicated support team
Kubernetes & Docker support
> Learn about Cloud On-Premises

Powerful, intuitive monitoring interface

Modern, responsive UI built for real-time troubleshooting with customizable dashboards and advanced visualization capabilities.

Real-time chart updates
Customizable dashboards
Dark & light themes
Advanced filtering & search
Responsive on all devices
Collaboration features
> Explore Netdata UI

Monitor on the go

Native iOS and Android apps bring full monitoring capabilities to your mobile device with real-time alerts and notifications.

iOS & Android apps
Push notifications
Touch-optimized interface
Offline data access
Biometric authentication
Widget support
> Download apps

The future of infrastructure observability

See our strategic direction across AI-native observability, full-stack signals, operational intelligence, and enterprise platform maturity.

AI-native observability
Full-stack signal coverage
Operational intelligence
Enterprise platform maturity
Agent releases every 6 weeks
Cloud continuous delivery
> Explore Product Roadmap

Best energy efficiency

True real-time per-second

100% automated zero config

Centralized observability

Multi-year retention

High availability built-in

Zero maintenance

Always up-to-date

Enterprise security

Complete data control

Air-gap ready

Compliance certified

Millisecond responsiveness

Infinite zoom & pan

Works on any device

Native performance

Instant alerts

Monitor anywhere

AI-native observability

Continuous delivery

Open source foundation

80% Faster Incident Resolution

AI-powered troubleshooting from detection, to root cause and blast radius identification, to reporting.

True Real-Time and Simple, even at Scale

Linearly and infinitely scalable full-stack observability, that can be deployed even mid-crisis.

90% Cost Reduction, Full Fidelity

Instead of centralizing the data, Netdata distributes the code, eliminating pipelines and complexity.

See and Map Your Entire Network

Live topology, flow analytics, and SNMP device and trap monitoring — unified with your full-stack observability.

Control Without Surrender

SOC 2 Type 2 certified with every metric kept on your infrastructure.

Integrations

800+ collectors and notification channels, auto-discovered and ready out of the box.

800+ data collectors
Auto-discovery & zero config
Cloud, infra, app protocols
Notifications out of the box
> Explore integrations
Real Results
46% Cost Reduction

Reduced monitoring costs by 46% while cutting staff overhead by 67%.

— Leonardo Antunez, Codyas

Zero Pipeline

No data shipping. No central storage costs. Query at the edge.

From Our Users
"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

No Query Language

Point-and-click troubleshooting. No PromQL, no LogQL, no learning curve.

Enterprise Ready
67% Less Staff, 46% Cost Cut

Enterprise efficiency without enterprise complexity—real ROI from day one.

— Leonardo Antunez, Codyas

SOC 2 Type 2 Certified

Zero data egress. Only metadata reaches the cloud. Your metrics stay on your infrastructure.

Full Coverage
800+ Collectors

Auto-discovered and configured. No manual setup required.

Any Notification Channel

Slack, PagerDuty, Teams, email, webhooks—all built-in.

Built for the People Who Get Paged

Because 3am alerts deserve instant answers, not hour-long hunts.

Every Industry Has Rules. We Master Them.

See how healthcare, finance, and government teams cut monitoring costs 90% while staying audit-ready.

Monitor Any Technology. Configure Nothing.

Install the agent. It already knows your stack.
From Our Users
"A Rare Unicorn"

Netdata gives more than you invest in it. A rare unicorn that obeys the Pareto rule.

— Eduard Porquet Mateu, TMB Barcelona

99% Downtime Reduction

Reduced website downtime by 99% and cloud bill by 30% using Netdata alerts.

— Falkland Islands Government

Real Savings
30% Cloud Cost Reduction

Optimized resource allocation based on Netdata alerts cut cloud spending by 30%.

— Falkland Islands Government

46% Cost Cut

Reduced monitoring staff by 67% while cutting operational costs by 46%.

— Codyas

Real Coverage
"Plugin for Everything"

Netdata has agent capacity or a plugin for everything, including Windows and Kubernetes.

— Eduard Porquet Mateu, TMB Barcelona

"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

Real Speed
Troubleshooting in 30 Seconds

From 2-3 minutes to 30 seconds—instant visibility into any node issue.

— Matthew Artist, Nodecraft

20% Downtime Reduction

20% less downtime and 40% budget optimization from out-of-the-box monitoring.

— Simon Beginn, LANCOM Systems

Pay per Node. Unlimited Everything Else.

One price per node. Unlimited metrics, logs, users, and retention. No per-GB surprises.

Free tier—forever
No metric limits or caps
Retention you control
Cancel anytime
> See pricing plans

What's Your Monitoring Really Costing You?

Most teams overpay by 40-60%. Let's find out why.

Expose hidden metric charges
Calculate tool consolidation
Customers report 30-67% savings
Results in under 60 seconds
> See what you're really paying

Your Infrastructure Is Unique. Let's Talk.

Because monitoring 10 nodes is different from monitoring 10,000.

On-prem & air-gapped deployment
Volume pricing & agreements
Architecture review for your scale
Compliance & security support
> Start a conversation

Monitoring That Sells Itself

Deploy in minutes. Impress clients in hours. Earn recurring revenue for years.

30-second live demos close deals
Zero config = zero support burden
Competitive margins & deal protection
Response in 48 hours
> Apply to partner

Per-Second Metrics at Homelab Prices

Same engine, same dashboards, same ML. Just priced for tinkerers.

Community: Free forever · 5 nodes · non-commercial
Homelab: $90/yr · unlimited nodes · fair usage
> Get the Homelab Plan

$1,000 Per Referral. Unlimited Referrals.

Your colleagues get 10% off. You get 10% commission. Everyone wins.

10% of subscriptions, up to $1,000 each
Track earnings inside Netdata Cloud
PayPal/Venmo payouts in 3-4 weeks
No caps, no complexity
> Get your referral link
Cost Proof
40% Budget Optimization

"Netdata's significant positive impact" — LANCOM Systems

Calculate Your Savings

Compare vs Datadog, Grafana, Dynatrace

Savings Proof
46% Cost Reduction

"Cut costs by 46%, staff by 67%" — Codyas

30% Cloud Bill Savings

"Reduced cloud bill by 30%" — Falkland Islands Gov

Enterprise Proof
"Better Than Combined Alternatives"

"Better observability with Netdata than combining other tools." — TMB Barcelona

Real Engineers, <24h Response

DPA, SLAs, on-prem, volume pricing

Why Partners Win
Demo Live Infrastructure

One command, 30 seconds, real data—no sandbox needed

Zero Tickets, High Margins

Auto-config + per-node pricing = predictable profit

Homelab Ready
Free Video Course

8-episode Netdata tutorial by LearnLinux.tv

76k+ GitHub Stars

3rd most starred monitoring project

Worth Recommending
Product That Delivers

Customers report 40-67% cost cuts, 99% downtime reduction

Zero Risk to Your Rep

Free tier lets them try before they buy

AI Support Assistant, Available 24/7

Nedi has access to all official documentation, source code, and resources. Ask any question about Netdata—responds in your language.

Deployment & configuration
Troubleshooting & sizing
Alerts & notifications
Evidence-based answers
> Ask Nedi now

Never Fight Fires Alone

Docs, community, and expert help—pick your path to resolution.

Learn.netdata.cloud docs
Discord, Forums, GitHub
Premium support available
> Get answers now

60 Seconds to First Dashboard

One command to install. Zero config. 850+ integrations documented.

Linux, Windows, K8s, Docker
Auto-discovers your stack
> Read our documentation

76,000+ Engineers Strong

615+ contributors. 1.5M daily downloads. One mission: simplify observability.

Per-Second. 90% Cheaper. Data Stays Home.

Side-by-side comparisons: costs, real-time granularity, and data sovereignty for every major tool.

See why teams switch from Datadog, Prometheus, Grafana, and more.

> Browse all comparisons
Edge-Native Observability, Born Open Source
Per-second visibility, ML on every metric, and data that never leaves your infrastructure.
Founded in 2016
615+ contributors worldwide
Remote-first, engineering-driven
Open source first
> Read our story
Promises We Publish—and Prove
12 principles backed by open code, independent validation, and measurable outcomes.
Open source, peer-reviewed
Zero config, instant value
Data sovereignty by design
Aligned pricing, no surprises
> See all 12 principles
Edge-Native, AI-Ready, 100% Open
76k+ stars. Full ML, AI, and automation—GPLv3+, not premium add-ons.
76,000+ GitHub stars
GPLv3+ licensed forever
ML on every metric, included
Zero vendor lock-in
> Explore our open source
Build Real-Time Observability for the World
Remote-first team shipping per-second monitoring with ML on every metric.
Remote-first, fully distributed
Open source (76k+ stars)
Challenging technical problems
Your code on millions of systems
> See open roles
Meet the Team Behind Netdata
Conferences, meetups, and tradeshows where you can see Netdata in action and talk to the engineers who build it.
Live demos and deep dives
Book 1-on-1 meetings
Talks and panel sessions
Event recaps and photos
> See all events
Talk to a Netdata Human in <24 Hours
Sales, partnerships, press, or professional services—real engineers, fast answers.
Discuss your observability needs
Pricing and volume discounts
Partnership opportunities
Media and press inquiries
> Book a conversation
Your Data. Your Rules.
On-prem data, cloud control plane, transparent terms.
Trust & Scale
76,000+ GitHub Stars

One of the most popular open-source monitoring projects

SOC 2 Type 2 Certified

Enterprise-grade security and compliance

Data Sovereignty

Your metrics stay on your infrastructure

Validated
University of Amsterdam

"Most energy-efficient monitoring solution" — ICSOC 2023, peer-reviewed

ADASTEC (Autonomous Driving)

"Doesn't miss alerts—mission-critical trust for safety software"

Community Stats
615+ Contributors

Global community improving monitoring for everyone

1.5M+ Downloads/Day

Trusted by teams worldwide

GPLv3+ Licensed

Free forever, fully open source agent

Why Join?
Remote-First

Work from anywhere, async-friendly culture

Impact at Scale

Your work helps millions of systems

$ guides / memcached / memcached-cpu-saturation ▌

Operations Guides

Memcached CPU saturation: the per-thread ceiling that aggregate graphs hide

Memcached is rarely CPU-bound. At normal request rates, worker threads spend most of their time in epoll waits. When CPU climbs on a memcached host, the reflex is to check aggregate rusage and conclude the daemon is busy. That reflex misses the actual failure mode.

Memcached dispatches each accepted connection round-robin to one of -t worker threads (default 4). Once assigned, a connection lives on that worker for its lifetime. The worker runs its own libevent loop and serves its assigned connections exclusively. If one worker saturates at 100% of a core, every connection pinned to that worker sees elevated latency. The other workers may be at 20%. The aggregate graph reads roughly 40%. Three quarters of your clients are fine.

Aggregate CPU, aggregate throughput, and even aggregate hit ratio can all look healthy while a quarter of your traffic suffers. The signal you need is per-thread, and memcached does not expose per-thread CPU directly. You derive it from OS-level thread views (top -H, ps -L, /proc/<pid>/task/<tid>/stat) and correlate it with conn_yields, which rises when workers cannot keep up with the request burst on their connections.

What this means

Two mechanisms cause the visible damage.

Per-thread ceiling. Each worker is single-threaded with no intra-thread parallelism. If the connections on worker 3 send requests faster than worker 3 can service them, requests queue inside that worker’s event loop. The clients on worker 3 see increasing latency. Clients on workers 0, 1, and 2 are unaffected.

conn_yields protection. When a single connection tries to run too many requests in one event loop iteration (default 20, set by -R), the worker yields that connection to the back of the queue and serves others. This is a fairness mechanism. A rising conn_yields rate means a worker is protecting itself from one aggressive connection at the cost of latency for that connection.

The two combine. If the workload saturates threads, conn_yields climbs with CPU. If only one client is misbehaving, conn_yields climbs without aggregate CPU saturation. Distinguishing the two is the diagnostic job.

A third CPU consumer runs in the background: the hash table expansion thread. When the cache fills and the hash table needs to grow, hash_is_expanding flips to 1 and the maintenance thread migrates items from the old table to the new one. This runs in a dedicated thread and does not block workers, but it does consume CPU and memory (the old and new tables coexist during migration). Correlate CPU spikes with hash_is_expanding = 1 before assuming request load is the cause.

flowchart TD
    A[Aggregate CPU climbing] --> B{Per-thread CPU balanced?}
    B -- No, one thread at 100% --> C[Per-thread ceiling]
    B -- Yes, all threads high --> D[Insufficient workers]
    B -- Yes, all threads low --> E[CPU spike not from workers]
    C --> F{conn_yields rising?}
    F -- Yes, with CPU --> G[Request volume exceeds capacity]
    F -- Yes, without CPU --> H[One aggressive client]
    F -- No --> I[Check hash_is_expanding]
    I --> J{hash_is_expanding = 1?}
    J -- Yes --> K[Background expansion]
    J -- No --> L[Check VmSwap, NIC, version]

Common causes

CauseWhat it looks likeFirst thing to check
Request rate exceeding per-thread capacityOne or two worker threads at 100%, others underutilized; conn_yields rising slowly; latency elevated for a subset of clientstop -H -p $(pgrep memcached) to find saturated threads
Aggressive single-client pipeliningconn_yields rising fast while aggregate CPU is moderate; one client dominatesClient-side metrics; -R value in startup flags
Hash table expansionhash_is_expanding = 1; CPU climbs on the maintenance thread; may correlate with curr_items crossing a power of 2stats for hash_is_expanding and hash_power_level
Pathological key patternsCPU rises without commensurate cmd_get increase; one worker stuck in lookuphash_power_level relative to curr_items; version check
Insufficient worker threadsAll worker threads at high utilization with idle cores available; throughput plateausStartup flags for -t; compare to host core count
Version-specific hash expansion bugOn memcached 1.5.14 through 1.6.13 at the maximum hashpower of 32, an integer overflow can hang the hash tableversion command; upgrade to 1.6.14+ if in the affected range

Quick checks

# Check daemon responsiveness and version
echo "version" | nc -q1 localhost 11211

# Cumulative CPU (per-process). Sample twice, 10s apart, compute delta.
echo "stats" | nc -q1 localhost 11211 | grep "STAT rusage"

# Connection fairness signal (cumulative). Derive rate across samples.
echo "stats" | nc -q1 localhost 11211 | grep "STAT conn_yields"

# Is the hash table expanding right now?
echo "stats" | nc -q1 localhost 11211 | grep -E "STAT hash_(is_expanding|power_level|bytes)"

# Per-thread CPU at the OS level (press H in top for thread view)
top -H -p $(pgrep memcached)

# Same data, scriptable. Sort by CPU to find the hot thread.
ps -L -o tid,pcpu,comm -p $(pgrep memcached) | sort -k2 -rn | head

# Startup flags: -t (threads), -R (requests per event), -c (maxconns)
ps -o args= -p $(pgrep memcached)

# Rule out swap (any nonzero value is a production incident)
grep VmSwap /proc/$(pgrep memcached)/status

The -q1 flag works with netcat-openbsd (Debian/Ubuntu default). On RHEL/CentOS with nmap-ncat, substitute -w1. Avoid polling stats more frequently than every 10 seconds under high load.

How to diagnose it

  1. Confirm aggregate CPU is genuinely climbing. Sample rusage_user and rusage_system twice, 10 seconds apart. The rate is the delta divided by the sample interval. A 4-thread memcached can use up to roughly 4 CPU-seconds per wall second.

  2. Check per-thread CPU. This is the step most operators skip. top -H -p $(pgrep memcached) or ps -L -o tid,pcpu,comm -p $(pgrep memcached) shows CPU per thread. If one thread is near 100% of a core and others are much lower, you have a per-thread ceiling problem, not a global CPU problem.

  3. Check conn_yields rate. Sample twice and compute the delta. If conn_yields is climbing with CPU, request volume is overwhelming at least one worker. If conn_yields is climbing but CPU is moderate, one client is sending large pipelines and the -R limit is doing its job.

  4. Check hash_is_expanding. If this is 1, the hash table is mid-expansion. Check hash_power_level and curr_items for context. Expansion should complete in seconds to minutes. If it persists, suspect a very large item count or a version-specific bug.

  5. Check version. Versions 1.5.14 through 1.6.13 may be affected by an integer overflow in the 32-bit hashsize calculation when hashpower reaches its maximum of 32; the result is a hang, not ordinary CPU saturation. 1.6.14 fixes it.

  6. Check per-client contribution if possible. Memcached does not expose per-connection command counts natively. Infer from client-side metrics, network-level observation, or by temporarily isolating suspect clients.

  7. Rule out adjacent causes. Swap (VmSwap in /proc/<pid>/status) causes CPU burn via page faults and is catastrophic for latency. NIC saturation (bytes_written versus link capacity) causes kernel system time to rise without user time rising proportionally.

Metrics and signals to monitor

SignalWhy it mattersWarning sign
rusage_user + rusage_system ratePer-process CPU consumption. Must be derived from deltas.Sustained rate approaching -t worker count
Per-thread CPU (OS-level)The only way to see the per-thread ceiling. Aggregate hides it entirely.Any single worker thread sustained near 100% of a core
conn_yields rateFairness mechanism activity. Confirms workers are throttling aggressive connections.Sustained non-zero rate, especially correlated with CPU
hash_is_expandingBackground maintenance CPU consumer.Persists beyond 60 seconds, or recurs frequently
hash_power_levelHash table size as power of 2. Context for expansion events.Reaching 32 on versions with the known overflow bug
cmd_get + cmd_set rateWorkload volume. Correlate against CPU rate.CPU rising faster than command rate suggests non-request overhead
bytes_written rateNetwork saturation drives kernel system time.Approaching NIC link capacity
VmSwap for memcachedSwap causes CPU-burning page faults and destroys latency.Any nonzero value
Client-observed latencyThe actual user impact signal. Not exposed by memcached natively.p99 climbing while p50 stays stable

Fixes

Per-thread saturation from request volume

Increase worker thread count. Bump -t at startup. The manpage recommends not exceeding the number of CPU cores, and 64 or more threads is not recommended. This requires a restart, which means total cache loss. Weigh against splitting traffic across more memcached instances, which also gives you more listener capacity. There is only one listener thread accepting new connections, and under heavy connection churn it can itself become a bottleneck.

Reduce per-request work. Larger values mean more serialization and more bytes to copy per response. If bytes_written / cmd_get is climbing, item size inflation is contributing to CPU. Compress large values on the client side before storage.

Upgrade to benefit from response batching. Pre-1.6.0 versions lack automatic batching of response syscalls. Upgrading can reduce server CPU by up to 25% when pipelining averages about 1.5 keys per syscall.

Aggressive single-client pipelining

Evaluate -R tuning. The default of 20 means a connection can send 20 requests per event loop iteration before being yielded to the back of the queue. If your workload legitimately pipelines more than 20 requests per event and the client is not starving others, raising -R reduces conn_yields overhead. If the client is starving others, do not raise -R. Fix the client instead. Changing -R requires a restart.

Fix the client. A single client sending tight-loop requests or massive unbatched multigets is an application-level bug. Identify it through client-side telemetry or network-level observation and rate-limit or batch at the application layer.

Hash table expansion

Let it finish. Expansion is normal and self-completing. If hash_is_expanding is 1 for a few seconds, do nothing. Monitor hash_power_level and curr_items to understand the growth trajectory.

Presize the hash table on large caches. For deployments with very large item counts, runtime expansion to high hashpower values consumes CPU and temporarily doubles hash table memory. Start the instance with -o hashpower=N (accepted range 12–32) to avoid runtime expansion.

Upgrade if affected by the overflow bug. The hashsize integer overflow applies to 1.5.14 through 1.6.13 when hashpower is 32. Memcached 1.6.14 fixes it.

Insufficient workers with idle cores

Increase -t to match the workload. If all worker threads are at high utilization and the host has idle cores, adding workers (up to core count) spreads connections across more event loops. Restart required, with total cache loss.

Prefer more instances over more threads at scale. Beyond a point, adding threads increases contention on the item lock table. Upstream sizes that table from 2^10 locks for fewer than three threads up to 2^15 locks for more than 20 threads. Multiple smaller instances on the same host, each with its own listener thread and slab allocator, can outperform one large instance with many threads.

Prevention

  • Monitor per-thread CPU, not just aggregate. Alert on any single memcached thread sustained above 80% of a core for 5 or more minutes. This catches the problem before clients do.
  • Track conn_yields rate as a capacity signal. A sustained non-zero rate means workers are actively throttling. Investigate before it becomes a latency incident.
  • Size -t to the host and workload. Match worker count to core count. Do not exceed it.
  • Keep memcached current. The hash expansion overflow fix and response batching are both meaningful for CPU behavior. Know which features your version has.
  • Monitor hash_is_expanding duration. Expansion persisting beyond 60 seconds or recurring frequently indicates item count instability.
  • Measure client-observed latency externally. Memcached exposes no latency histograms natively. Without client-side measurement, you cannot see the per-thread ceiling’s impact on users.

How Netdata helps

  • Netdata collects rusage_user and rusage_system per second and derives rates automatically, so you see current CPU consumption without manual delta math.
  • Per-thread CPU from the OS level is collected alongside the memcached stats context, letting you correlate aggregate rusage with actual per-thread saturation in the same time window.
  • conn_yields rate is tracked alongside cmd_get and cmd_set, so the relationship between request volume, fairness throttling, and CPU is visible on one chart.
  • Correlated dashboards let you overlay memcached CPU with NIC utilization and VmSwap, so you can rule out network saturation and swap as the cause.