The only agent that thinks for itself

Autonomous Monitoring with self-learning AI built-in, operating independently across your entire stack.

Unlimited Metrics & Logs
Machine learning & MCP
5% CPU, 150MB RAM
3GB disk, >1 year retention
800+ integrations, zero config
Dashboards, alerts out of the box
> Discover Netdata Agents

Centralized metrics streaming and storage

Aggregate metrics from multiple agents into centralized Parent nodes for unified monitoring across your infrastructure.

Stream from unlimited agents
Long-term data retention
High availability clustering
Data replication & backup
Scalable architecture
Enterprise-grade security
> Learn about Parents

Fully managed cloud platform

Access your monitoring data from anywhere with our SaaS platform. No infrastructure to manage, automatic updates, and global availability.

Zero infrastructure management
99.9% uptime SLA
Global data centers
Automatic updates & patches
Enterprise SSO & RBAC
SOC2 & ISO certified
> Explore Netdata Cloud

Deploy Netdata Cloud in your infrastructure

Run the full Netdata Cloud platform on-premises for complete data sovereignty and compliance with your security policies.

Complete data sovereignty
Air-gapped deployment
Custom compliance controls
Private network integration
Dedicated support team
Kubernetes & Docker support
> Learn about Cloud On-Premises

Powerful, intuitive monitoring interface

Modern, responsive UI built for real-time troubleshooting with customizable dashboards and advanced visualization capabilities.

Real-time chart updates
Customizable dashboards
Dark & light themes
Advanced filtering & search
Responsive on all devices
Collaboration features
> Explore Netdata UI

Monitor on the go

Native iOS and Android apps bring full monitoring capabilities to your mobile device with real-time alerts and notifications.

iOS & Android apps
Push notifications
Touch-optimized interface
Offline data access
Biometric authentication
Widget support
> Download apps

The future of infrastructure observability

See our strategic direction across AI-native observability, full-stack signals, operational intelligence, and enterprise platform maturity.

AI-native observability
Full-stack signal coverage
Operational intelligence
Enterprise platform maturity
Agent releases every 6 weeks
Cloud continuous delivery
> Explore Product Roadmap

Best energy efficiency

True real-time per-second

100% automated zero config

Centralized observability

Multi-year retention

High availability built-in

Zero maintenance

Always up-to-date

Enterprise security

Complete data control

Air-gap ready

Compliance certified

Millisecond responsiveness

Infinite zoom & pan

Works on any device

Native performance

Instant alerts

Monitor anywhere

AI-native observability

Continuous delivery

Open source foundation

80% Faster Incident Resolution

AI-powered troubleshooting from detection, to root cause and blast radius identification, to reporting.

True Real-Time and Simple, even at Scale

Linearly and infinitely scalable full-stack observability, that can be deployed even mid-crisis.

90% Cost Reduction, Full Fidelity

Instead of centralizing the data, Netdata distributes the code, eliminating pipelines and complexity.

See and Map Your Entire Network

Live topology, flow analytics, and SNMP device and trap monitoring — unified with your full-stack observability.

Control Without Surrender

SOC 2 Type 2 certified with every metric kept on your infrastructure.

Integrations

800+ collectors and notification channels, auto-discovered and ready out of the box.

800+ data collectors
Auto-discovery & zero config
Cloud, infra, app protocols
Notifications out of the box
> Explore integrations
Real Results
46% Cost Reduction

Reduced monitoring costs by 46% while cutting staff overhead by 67%.

— Leonardo Antunez, Codyas

Zero Pipeline

No data shipping. No central storage costs. Query at the edge.

From Our Users
"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

No Query Language

Point-and-click troubleshooting. No PromQL, no LogQL, no learning curve.

Enterprise Ready
67% Less Staff, 46% Cost Cut

Enterprise efficiency without enterprise complexity—real ROI from day one.

— Leonardo Antunez, Codyas

SOC 2 Type 2 Certified

Zero data egress. Only metadata reaches the cloud. Your metrics stay on your infrastructure.

Full Coverage
800+ Collectors

Auto-discovered and configured. No manual setup required.

Any Notification Channel

Slack, PagerDuty, Teams, email, webhooks—all built-in.

Built for the People Who Get Paged

Because 3am alerts deserve instant answers, not hour-long hunts.

Every Industry Has Rules. We Master Them.

See how healthcare, finance, and government teams cut monitoring costs 90% while staying audit-ready.

Monitor Any Technology. Configure Nothing.

Install the agent. It already knows your stack.
From Our Users
"A Rare Unicorn"

Netdata gives more than you invest in it. A rare unicorn that obeys the Pareto rule.

— Eduard Porquet Mateu, TMB Barcelona

99% Downtime Reduction

Reduced website downtime by 99% and cloud bill by 30% using Netdata alerts.

— Falkland Islands Government

Real Savings
30% Cloud Cost Reduction

Optimized resource allocation based on Netdata alerts cut cloud spending by 30%.

— Falkland Islands Government

46% Cost Cut

Reduced monitoring staff by 67% while cutting operational costs by 46%.

— Codyas

Real Coverage
"Plugin for Everything"

Netdata has agent capacity or a plugin for everything, including Windows and Kubernetes.

— Eduard Porquet Mateu, TMB Barcelona

"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

Real Speed
Troubleshooting in 30 Seconds

From 2-3 minutes to 30 seconds—instant visibility into any node issue.

— Matthew Artist, Nodecraft

20% Downtime Reduction

20% less downtime and 40% budget optimization from out-of-the-box monitoring.

— Simon Beginn, LANCOM Systems

Pay per Node. Unlimited Everything Else.

One price per node. Unlimited metrics, logs, users, and retention. No per-GB surprises.

Free tier—forever
No metric limits or caps
Retention you control
Cancel anytime
> See pricing plans

What's Your Monitoring Really Costing You?

Most teams overpay by 40-60%. Let's find out why.

Expose hidden metric charges
Calculate tool consolidation
Customers report 30-67% savings
Results in under 60 seconds
> See what you're really paying

Your Infrastructure Is Unique. Let's Talk.

Because monitoring 10 nodes is different from monitoring 10,000.

On-prem & air-gapped deployment
Volume pricing & agreements
Architecture review for your scale
Compliance & security support
> Start a conversation

Monitoring That Sells Itself

Deploy in minutes. Impress clients in hours. Earn recurring revenue for years.

30-second live demos close deals
Zero config = zero support burden
Competitive margins & deal protection
Response in 48 hours
> Apply to partner

Per-Second Metrics at Homelab Prices

Same engine, same dashboards, same ML. Just priced for tinkerers.

Community: Free forever · 5 nodes · non-commercial
Homelab: $90/yr · unlimited nodes · fair usage
> Get the Homelab Plan

$1,000 Per Referral. Unlimited Referrals.

Your colleagues get 10% off. You get 10% commission. Everyone wins.

10% of subscriptions, up to $1,000 each
Track earnings inside Netdata Cloud
PayPal/Venmo payouts in 3-4 weeks
No caps, no complexity
> Get your referral link
Cost Proof
40% Budget Optimization

"Netdata's significant positive impact" — LANCOM Systems

Calculate Your Savings

Compare vs Datadog, Grafana, Dynatrace

Savings Proof
46% Cost Reduction

"Cut costs by 46%, staff by 67%" — Codyas

30% Cloud Bill Savings

"Reduced cloud bill by 30%" — Falkland Islands Gov

Enterprise Proof
"Better Than Combined Alternatives"

"Better observability with Netdata than combining other tools." — TMB Barcelona

Real Engineers, <24h Response

DPA, SLAs, on-prem, volume pricing

Why Partners Win
Demo Live Infrastructure

One command, 30 seconds, real data—no sandbox needed

Zero Tickets, High Margins

Auto-config + per-node pricing = predictable profit

Homelab Ready
Free Video Course

8-episode Netdata tutorial by LearnLinux.tv

76k+ GitHub Stars

3rd most starred monitoring project

Worth Recommending
Product That Delivers

Customers report 40-67% cost cuts, 99% downtime reduction

Zero Risk to Your Rep

Free tier lets them try before they buy

AI Support Assistant, Available 24/7

Nedi has access to all official documentation, source code, and resources. Ask any question about Netdata—responds in your language.

Deployment & configuration
Troubleshooting & sizing
Alerts & notifications
Evidence-based answers
> Ask Nedi now

Never Fight Fires Alone

Docs, community, and expert help—pick your path to resolution.

Learn.netdata.cloud docs
Discord, Forums, GitHub
Premium support available
> Get answers now

60 Seconds to First Dashboard

One command to install. Zero config. 850+ integrations documented.

Linux, Windows, K8s, Docker
Auto-discovers your stack
> Read our documentation

76,000+ Engineers Strong

615+ contributors. 1.5M daily downloads. One mission: simplify observability.

Per-Second. 90% Cheaper. Data Stays Home.

Side-by-side comparisons: costs, real-time granularity, and data sovereignty for every major tool.

See why teams switch from Datadog, Prometheus, Grafana, and more.

> Browse all comparisons
Edge-Native Observability, Born Open Source
Per-second visibility, ML on every metric, and data that never leaves your infrastructure.
Founded in 2016
615+ contributors worldwide
Remote-first, engineering-driven
Open source first
> Read our story
Promises We Publish—and Prove
12 principles backed by open code, independent validation, and measurable outcomes.
Open source, peer-reviewed
Zero config, instant value
Data sovereignty by design
Aligned pricing, no surprises
> See all 12 principles
Edge-Native, AI-Ready, 100% Open
76k+ stars. Full ML, AI, and automation—GPLv3+, not premium add-ons.
76,000+ GitHub stars
GPLv3+ licensed forever
ML on every metric, included
Zero vendor lock-in
> Explore our open source
Build Real-Time Observability for the World
Remote-first team shipping per-second monitoring with ML on every metric.
Remote-first, fully distributed
Open source (76k+ stars)
Challenging technical problems
Your code on millions of systems
> See open roles
Meet the Team Behind Netdata
Conferences, meetups, and tradeshows where you can see Netdata in action and talk to the engineers who build it.
Live demos and deep dives
Book 1-on-1 meetings
Talks and panel sessions
Event recaps and photos
> See all events
Talk to a Netdata Human in <24 Hours
Sales, partnerships, press, or professional services—real engineers, fast answers.
Discuss your observability needs
Pricing and volume discounts
Partnership opportunities
Media and press inquiries
> Book a conversation
Your Data. Your Rules.
On-prem data, cloud control plane, transparent terms.
Trust & Scale
76,000+ GitHub Stars

One of the most popular open-source monitoring projects

SOC 2 Type 2 Certified

Enterprise-grade security and compliance

Data Sovereignty

Your metrics stay on your infrastructure

Validated
University of Amsterdam

"Most energy-efficient monitoring solution" — ICSOC 2023, peer-reviewed

ADASTEC (Autonomous Driving)

"Doesn't miss alerts—mission-critical trust for safety software"

Community Stats
615+ Contributors

Global community improving monitoring for everyone

1.5M+ Downloads/Day

Trusted by teams worldwide

GPLv3+ Licensed

Free forever, fully open source agent

Why Join?
Remote-First

Work from anywhere, async-friendly culture

Impact at Scale

Your work helps millions of systems

$ guides / memcached / memcached-network-saturation ▌

Operations Guides

Memcached network saturation: large values, multiget, and a maxed-out NIC

Memcached latency is climbing for every client, but the host looks healthy. CPU is well below saturation, memory utilization is far from the limit, evictions are zero, and curr_connections is well below -c. The daemon answers version and stats instantly. Yet every application reports slow gets, and the slowdown is uniform across keys and clients.

When latency degrades uniformly and the usual suspects are clear, look at the wire. Memcached pulls values from RAM and writes them to a socket. If the response stream exceeds what the NIC can push, the kernel queues, TCP congestion control engages, and every client waits. The cache is fine; the link is the bottleneck.

What this means

Network bandwidth saturation occurs when bytes_written (the cumulative bytes the daemon has pushed to the network) approaches NIC capacity and the kernel cannot drain the send queue as fast as memcached produces responses.

The problem is asymmetric. Memcached responses carry values; requests carry keys. For read-heavy workloads, bytes_written is typically much larger than bytes_read. A workload storing multi-KB values and reading them at even a moderate rate can push more bytes per second than the link can carry, while operations per second remains well below the daemon’s CPU ceiling.

Multiget amplifies the effect. A single multiget for 100 keys at 10 KB each produces a 1 MB response assembled and written as one logical operation. The command rate barely moves, but bytes_written spikes. A few thousand such multigets per second can saturate a 1 Gbps link.

The signature is uniform latency degradation. Every client, key, and operation slows down together. CPU and memory remain healthy. The only signals that move are the bytes_written rate and OS NIC TX counters. Past roughly 85% of link speed, packet drops and retransmissions begin, and latency spikes.

flowchart TD
    A[Large values stored in cache] --> B[High get rate or large multigets]
    B --> C[bytes_written rate climbs]
    C --> D[NIC TX approaches link speed]
    D --> E{Share of link speed}
    E -->|> 70% sustained| F[Latency rises uniformly]
    E -->|> 85%| G[Packet drops and retransmits]
    F --> H[All clients slow down together]
    G --> H
    H --> I[CPU and memory still look fine]

These are application-layer bytes. bytes_written excludes TCP/IP overhead, so actual NIC utilization runs 5-10% higher than the memcached counter suggests. Always confirm against OS-level NIC counters before concluding the link is the limit.

Common causes

CauseWhat it looks likeFirst thing to check
Large stored valuesHigh bytes_written / cmd_get ratio; items in multi-KB or MB rangestats slabs chunk sizes and stats items per-class counts
Large multiget batchescmd_get moderate but bytes_written spikes; clients issue gets with many keysClient-side multiget batch sizes and value sizes
Traffic surge to large valuesNew feature, campaign, or scraper reading big objectsApplication telemetry for per-key or per-client access
NIC too slow for workloadSustained bytes_written near link speed with healthy ops/secip -s link show TX bytes versus link speed
NIC misconfigurationHalf-duplex negotiation, offloads disabled, single IRQ pinnedethtool link mode, ethtool -k, /proc/interrupts

Quick checks

All commands below are read-only and safe to run in production.

# Cumulative bytes the daemon has read and written
echo "stats" | nc -q1 localhost 11211 | grep -E "STAT bytes_(read|written)"

# Command counters for computing average response size
echo "stats" | nc -q1 localhost 11211 | grep -E "STAT cmd_(get|set)"

# OS NIC byte and error counters
cat /proc/net/dev

# Per-interface TX bytes, packets, drops, overruns
ip -s link show <interface>

# NIC link speed and duplex mode
ethtool <interface>

# Offload feature state (look for gro, lro, tso)
ethtool -k <interface>

# IRQ distribution for the NIC across cores
grep <interface> /proc/interrupts

# Slab class chunk sizes to see where large values land
echo "stats slabs" | nc -q1 localhost 11211 | grep chunk_size

Use -w 2 if -q1 is unavailable for nc (for example, on macOS or certain BSD-based netcat ports).

How to diagnose it

  1. Confirm the symptom is uniform. Check whether latency rose across all clients and keys at the same time. Uniform degradation points to the server or network; per-client or per-key spikes point to the application. Memcached exposes no native latency histograms, so confirm this via client-side instrumentation.
  2. Compute the bytes_written rate. Sample bytes_written twice with a known interval. Subtract and divide by the interval to get bytes per second. A single sample is cumulative since process start and useless for rate calculation.
  3. Express the rate as a percentage of link speed. For a 1 Gbps link, full-duplex capacity is roughly 125 MB/s in one direction. Above 70% sustained (about 87 MB/s) is concerning; above 85% (about 106 MB/s), expect latency spikes and packet loss. Adjust the denominator for your actual link speed (10 Gbps is roughly 1.25 GB/s per direction).
  4. Confirm against OS NIC counters. Pull TX bytes from /proc/net/dev or ip -s link show. The kernel counter includes all interface traffic. If the interface is near saturation but memcached’s bytes_written rate is modest, another process is using the NIC. If both agree, memcached is the dominant producer.
  5. Check for drops and overruns. In the ip -s link output, watch TX drops and overruns. Any nonzero rate means the kernel ran out of ring buffer space. This is where latency stops climbing gradually and starts spiking.
  6. Determine whether the cause is value size or batch size. Compute bytes_written / cmd_get over the same window. Memcached increments cmd_get per requested key, not per command, so this ratio accurately reflects the average response size per key. A rising ratio over days or weeks means values are inflating. A sudden jump means either new large values or larger multiget batches entered the workload.
  7. Rule out NIC misconfiguration. Run ethtool <interface> to confirm the link negotiated at full duplex and expected speed. Half-duplex links cause throughput collapse. Run ethtool -k <interface> to confirm segmentation offloads (TSO, GRO, LRO) are enabled. Disabled offloads push per-packet processing to the CPU and reduce effective throughput.
  8. Check IRQ distribution. If all NIC interrupts land on one core, that core bottlenecks before the link does. Inspect /proc/interrupts for the NIC’s queues. Uneven distribution combined with high rusage_system on one core suggests the kernel network stack (softirq or send processing) is the constraint, not the link itself. rusage_system is the memcached process’s system CPU: its own syscalls, including inline TCP send processing. On kernels with IRQ time accounting (CONFIG_IRQ_TIME_ACCOUNTING, enabled on most current distros), softirq processing is accounted system-wide instead — visible as %soft in mpstat or as ksoftirqd usage — so check both before concluding.

Metrics and signals to monitor

SignalWhy it mattersWarning sign
bytes_written rateDirect measure of egress volume from the daemonSustained above 70% of link speed
bytes_written / cmd_getAverage response size per key; rising trend means value inflationSteady upward drift over days
bytes_read rateConfirms the read-heavy asymmetryFar smaller than bytes_written
OS NIC TX bytesGround truth from the kernel interfaceDisagrees with bytes_written means other traffic shares the NIC
NIC TX drops and overrunsKernel ring buffer exhaustionAny nonzero sustained rate
Client-observed p99 latencyThe actual user impact signalAbove 5 ms for same-datacenter memcached
cmd_get rateConfirms ops/sec is not the bottleneckModerate while bytes_written is near link speed
rusage_system rateSystem CPU time of the memcached processOne core pegged indicates kernel network processing limits

Memcached does not expose per-operation latency. The latency signal must come from the client. Without client instrumentation, the server-side signals above are your only diagnostic path.

Fixes

Reduce value size with client-side compression

Most memcached client libraries support transparent compression. The client compresses before set, sets a flag bit, and decompresses on retrieval. Text and binary payloads often compress 3x to 10x, which directly reduces bytes_written for the same cmd_get rate.

Compression adds CPU cost on the client, not the server. If your application servers have CPU headroom and the memcached host is network-bound, this is usually the highest-leverage fix. Measure compression ratios on your actual data before assuming a specific gain.

Shrink multiget batches

If clients issue multigets with dozens or hundreds of keys, splitting them into smaller batches reduces peak response size per operation. This does not reduce total bytes over time, but it smooths bursty transmission and gives the kernel more chances to drain the queue between responses.

This increases round trips per logical operation. It helps when the problem is burst-induced queue buildup, not sustained average throughput. If sustained throughput is the limit, smaller batches will not help.

Spread load across more instances

Adding memcached instances on different hosts distributes egress across multiple NICs. In a client-side sharded cluster, consistent hashing redistributes keys across the new nodes automatically. This is the right structural fix when a single instance is permanently over its link budget.

More hosts mean more operational overhead. Adding instances to the same host does not help because the bottleneck is the host’s NIC, not the memcached process. Sharding across more instances on one box only moves the queue from per-socket buffers to the kernel qdisc layer.

Upgrade the NIC

If the workload is legitimately large-value and the application cannot compress further, a faster link is the direct fix. Moving from 1 Gbps to 10 Gbps gives a 10x ceiling. At 400,000 operations per second with an average value size of 3072 bytes, even a 10 Gbps link saturates.

Confirm the kernel, PCIe, and interrupt handling can actually deliver the new link rate before paying for it. A 25 Gbps NIC on a host that cannot process interrupts fast enough buys nothing.

Verify NIC offloads and IRQ affinity

Before buying hardware, confirm the existing NIC is configured correctly. Distribute NIC receive and transmit interrupts across multiple cores. Misconfiguration can halve effective throughput.

Warning: ethtool -K changes live NIC settings and can cause a brief link flap or packet loss depending on the driver. Test during a maintenance window, not during an active incident.

# WARNING: disruptive on some drivers; verify current state with ethtool -k first
ethtool -K <interface> tso on gro on lro on

TSO is the most relevant offload for egress saturation. GRO and LRO are receive-side offloads. Offloads are the default on most distributions, but verify rather than assume. Some virtualization platforms and cloud providers expose limited offload control.

Prevention

  • Track bytes_written as a percentage of link speed. Alert on sustained utilization above 70%. Express it as a rate, not a cumulative counter.
  • Track the bytes_written / cmd_get ratio over time. A rising trend means average value size is inflating, even before the link saturates. Catch this in capacity planning, not during an incident.
  • Capacity plan against link speed, not just operations per second. Memcached handles hundreds of thousands of operations per second on modern hardware, but a few thousand large-value multigets can saturate a 1 Gbps link. Ops/sec and bytes/sec are independent ceilings.
  • Audit value sizes at the application layer. Serialized objects that grow over time (new fields, larger nested structures) silently push the workload toward the link limit. Review the largest cached objects during every major application release.
  • Confirm NIC configuration after host provisioning. Duplex mismatches and disabled offloads are provisioning-time errors that persist until someone investigates a slowdown. Build ethtool checks into your host bootstrap.

How Netdata helps

  • Per-second bytes_written and bytes_read collection from the memcached stats endpoint, with rates computed automatically. Per-second resolution catches burst-induced saturation that 60-second polling misses.
  • OS NIC TX metrics per interface, collected at the same per-second cadence. Correlating memcached bytes_written with kernel TX bytes in one view confirms whether the link is the limit or whether another process shares the NIC.
  • Derived bytes_written / cmd_get tracking as a computed dimension, so average response size trends are visible without manual sampling.
  • NIC drop and overrun counters surfaced alongside throughput, so the transition from gradual latency rise to packet loss is visible as it happens.
  • Anomaly detection on the bytes_written rate and NIC TX rate, flagging deviations from the learned baseline.
  • Client-side latency correlation when application instrumentation is available, confirming the uniform-degradation signature that distinguishes network saturation from CPU or memory pressure.