The only agent that thinks for itself

Autonomous Monitoring with self-learning AI built-in, operating independently across your entire stack.

Unlimited Metrics & Logs
Machine learning & MCP
5% CPU, 150MB RAM
3GB disk, >1 year retention
800+ integrations, zero config
Dashboards, alerts out of the box
> Discover Netdata Agents

Centralized metrics streaming and storage

Aggregate metrics from multiple agents into centralized Parent nodes for unified monitoring across your infrastructure.

Stream from unlimited agents
Long-term data retention
High availability clustering
Data replication & backup
Scalable architecture
Enterprise-grade security
> Learn about Parents

Fully managed cloud platform

Access your monitoring data from anywhere with our SaaS platform. No infrastructure to manage, automatic updates, and global availability.

Zero infrastructure management
99.9% uptime SLA
Global data centers
Automatic updates & patches
Enterprise SSO & RBAC
SOC2 & ISO certified
> Explore Netdata Cloud

Deploy Netdata Cloud in your infrastructure

Run the full Netdata Cloud platform on-premises for complete data sovereignty and compliance with your security policies.

Complete data sovereignty
Air-gapped deployment
Custom compliance controls
Private network integration
Dedicated support team
Kubernetes & Docker support
> Learn about Cloud On-Premises

Powerful, intuitive monitoring interface

Modern, responsive UI built for real-time troubleshooting with customizable dashboards and advanced visualization capabilities.

Real-time chart updates
Customizable dashboards
Dark & light themes
Advanced filtering & search
Responsive on all devices
Collaboration features
> Explore Netdata UI

Monitor on the go

Native iOS and Android apps bring full monitoring capabilities to your mobile device with real-time alerts and notifications.

iOS & Android apps
Push notifications
Touch-optimized interface
Offline data access
Biometric authentication
Widget support
> Download apps

The future of infrastructure observability

See our strategic direction across AI-native observability, full-stack signals, operational intelligence, and enterprise platform maturity.

AI-native observability
Full-stack signal coverage
Operational intelligence
Enterprise platform maturity
Agent releases every 6 weeks
Cloud continuous delivery
> Explore Product Roadmap

Best energy efficiency

True real-time per-second

100% automated zero config

Centralized observability

Multi-year retention

High availability built-in

Zero maintenance

Always up-to-date

Enterprise security

Complete data control

Air-gap ready

Compliance certified

Millisecond responsiveness

Infinite zoom & pan

Works on any device

Native performance

Instant alerts

Monitor anywhere

AI-native observability

Continuous delivery

Open source foundation

80% Faster Incident Resolution

AI-powered troubleshooting from detection, to root cause and blast radius identification, to reporting.

True Real-Time and Simple, even at Scale

Linearly and infinitely scalable full-stack observability, that can be deployed even mid-crisis.

90% Cost Reduction, Full Fidelity

Instead of centralizing the data, Netdata distributes the code, eliminating pipelines and complexity.

See and Map Your Entire Network

Live topology, flow analytics, and SNMP device and trap monitoring — unified with your full-stack observability.

Control Without Surrender

SOC 2 Type 2 certified with every metric kept on your infrastructure.

Integrations

800+ collectors and notification channels, auto-discovered and ready out of the box.

800+ data collectors
Auto-discovery & zero config
Cloud, infra, app protocols
Notifications out of the box
> Explore integrations
Real Results
46% Cost Reduction

Reduced monitoring costs by 46% while cutting staff overhead by 67%.

— Leonardo Antunez, Codyas

Zero Pipeline

No data shipping. No central storage costs. Query at the edge.

From Our Users
"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

No Query Language

Point-and-click troubleshooting. No PromQL, no LogQL, no learning curve.

Enterprise Ready
67% Less Staff, 46% Cost Cut

Enterprise efficiency without enterprise complexity—real ROI from day one.

— Leonardo Antunez, Codyas

SOC 2 Type 2 Certified

Zero data egress. Only metadata reaches the cloud. Your metrics stay on your infrastructure.

Full Coverage
800+ Collectors

Auto-discovered and configured. No manual setup required.

Any Notification Channel

Slack, PagerDuty, Teams, email, webhooks—all built-in.

Built for the People Who Get Paged

Because 3am alerts deserve instant answers, not hour-long hunts.

Every Industry Has Rules. We Master Them.

See how healthcare, finance, and government teams cut monitoring costs 90% while staying audit-ready.

Monitor Any Technology. Configure Nothing.

Install the agent. It already knows your stack.
From Our Users
"A Rare Unicorn"

Netdata gives more than you invest in it. A rare unicorn that obeys the Pareto rule.

— Eduard Porquet Mateu, TMB Barcelona

99% Downtime Reduction

Reduced website downtime by 99% and cloud bill by 30% using Netdata alerts.

— Falkland Islands Government

Real Savings
30% Cloud Cost Reduction

Optimized resource allocation based on Netdata alerts cut cloud spending by 30%.

— Falkland Islands Government

46% Cost Cut

Reduced monitoring staff by 67% while cutting operational costs by 46%.

— Codyas

Real Coverage
"Plugin for Everything"

Netdata has agent capacity or a plugin for everything, including Windows and Kubernetes.

— Eduard Porquet Mateu, TMB Barcelona

"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

Real Speed
Troubleshooting in 30 Seconds

From 2-3 minutes to 30 seconds—instant visibility into any node issue.

— Matthew Artist, Nodecraft

20% Downtime Reduction

20% less downtime and 40% budget optimization from out-of-the-box monitoring.

— Simon Beginn, LANCOM Systems

Pay per Node. Unlimited Everything Else.

One price per node. Unlimited metrics, logs, users, and retention. No per-GB surprises.

Free tier—forever
No metric limits or caps
Retention you control
Cancel anytime
> See pricing plans

What's Your Monitoring Really Costing You?

Most teams overpay by 40-60%. Let's find out why.

Expose hidden metric charges
Calculate tool consolidation
Customers report 30-67% savings
Results in under 60 seconds
> See what you're really paying

Your Infrastructure Is Unique. Let's Talk.

Because monitoring 10 nodes is different from monitoring 10,000.

On-prem & air-gapped deployment
Volume pricing & agreements
Architecture review for your scale
Compliance & security support
> Start a conversation

Monitoring That Sells Itself

Deploy in minutes. Impress clients in hours. Earn recurring revenue for years.

30-second live demos close deals
Zero config = zero support burden
Competitive margins & deal protection
Response in 48 hours
> Apply to partner

Per-Second Metrics at Homelab Prices

Same engine, same dashboards, same ML. Just priced for tinkerers.

Community: Free forever · 5 nodes · non-commercial
Homelab: $90/yr · unlimited nodes · fair usage
> Get the Homelab Plan

$1,000 Per Referral. Unlimited Referrals.

Your colleagues get 10% off. You get 10% commission. Everyone wins.

10% of subscriptions, up to $1,000 each
Track earnings inside Netdata Cloud
PayPal/Venmo payouts in 3-4 weeks
No caps, no complexity
> Get your referral link
Cost Proof
40% Budget Optimization

"Netdata's significant positive impact" — LANCOM Systems

Calculate Your Savings

Compare vs Datadog, Grafana, Dynatrace

Savings Proof
46% Cost Reduction

"Cut costs by 46%, staff by 67%" — Codyas

30% Cloud Bill Savings

"Reduced cloud bill by 30%" — Falkland Islands Gov

Enterprise Proof
"Better Than Combined Alternatives"

"Better observability with Netdata than combining other tools." — TMB Barcelona

Real Engineers, <24h Response

DPA, SLAs, on-prem, volume pricing

Why Partners Win
Demo Live Infrastructure

One command, 30 seconds, real data—no sandbox needed

Zero Tickets, High Margins

Auto-config + per-node pricing = predictable profit

Homelab Ready
Free Video Course

8-episode Netdata tutorial by LearnLinux.tv

76k+ GitHub Stars

3rd most starred monitoring project

Worth Recommending
Product That Delivers

Customers report 40-67% cost cuts, 99% downtime reduction

Zero Risk to Your Rep

Free tier lets them try before they buy

AI Support Assistant, Available 24/7

Nedi has access to all official documentation, source code, and resources. Ask any question about Netdata—responds in your language.

Deployment & configuration
Troubleshooting & sizing
Alerts & notifications
Evidence-based answers
> Ask Nedi now

Never Fight Fires Alone

Docs, community, and expert help—pick your path to resolution.

Learn.netdata.cloud docs
Discord, Forums, GitHub
Premium support available
> Get answers now

60 Seconds to First Dashboard

One command to install. Zero config. 850+ integrations documented.

Linux, Windows, K8s, Docker
Auto-discovers your stack
> Read our documentation

76,000+ Engineers Strong

615+ contributors. 1.5M daily downloads. One mission: simplify observability.

Per-Second. 90% Cheaper. Data Stays Home.

Side-by-side comparisons: costs, real-time granularity, and data sovereignty for every major tool.

See why teams switch from Datadog, Prometheus, Grafana, and more.

> Browse all comparisons
Edge-Native Observability, Born Open Source
Per-second visibility, ML on every metric, and data that never leaves your infrastructure.
Founded in 2016
615+ contributors worldwide
Remote-first, engineering-driven
Open source first
> Read our story
Promises We Publish—and Prove
12 principles backed by open code, independent validation, and measurable outcomes.
Open source, peer-reviewed
Zero config, instant value
Data sovereignty by design
Aligned pricing, no surprises
> See all 12 principles
Edge-Native, AI-Ready, 100% Open
76k+ stars. Full ML, AI, and automation—GPLv3+, not premium add-ons.
76,000+ GitHub stars
GPLv3+ licensed forever
ML on every metric, included
Zero vendor lock-in
> Explore our open source
Build Real-Time Observability for the World
Remote-first team shipping per-second monitoring with ML on every metric.
Remote-first, fully distributed
Open source (76k+ stars)
Challenging technical problems
Your code on millions of systems
> See open roles
Meet the Team Behind Netdata
Conferences, meetups, and tradeshows where you can see Netdata in action and talk to the engineers who build it.
Live demos and deep dives
Book 1-on-1 meetings
Talks and panel sessions
Event recaps and photos
> See all events
Talk to a Netdata Human in <24 Hours
Sales, partnerships, press, or professional services—real engineers, fast answers.
Discuss your observability needs
Pricing and volume discounts
Partnership opportunities
Media and press inquiries
> Book a conversation
Your Data. Your Rules.
On-prem data, cloud control plane, transparent terms.
Trust & Scale
76,000+ GitHub Stars

One of the most popular open-source monitoring projects

SOC 2 Type 2 Certified

Enterprise-grade security and compliance

Data Sovereignty

Your metrics stay on your infrastructure

Validated
University of Amsterdam

"Most energy-efficient monitoring solution" — ICSOC 2023, peer-reviewed

ADASTEC (Autonomous Driving)

"Doesn't miss alerts—mission-critical trust for safety software"

Community Stats
615+ Contributors

Global community improving monitoring for everyone

1.5M+ Downloads/Day

Trusted by teams worldwide

GPLv3+ Licensed

Free forever, fully open source agent

Why Join?
Remote-First

Work from anywhere, async-friendly culture

Impact at Scale

Your work helps millions of systems

$ guides / memcached / memcached-how-it-works-in-production ▌

Operations Guides

How Memcached actually works in production: a mental model for operators

Memcached is a multi-threaded, in-memory key-value cache daemon built on libevent. By default it has no persistence; it also has no replication or clustering. Clients handle sharding. A restart means total data loss. These are the design, not limitations to work around.

Before you can debug a memcached incident, you need three abstractions: how memory is partitioned (the slab allocator), how items age within each partition (the segmented LRU), and how connections and threads interact under load. Without these, the stats output is noise. With them, the same numbers tell you exactly which resource is saturated.

Daemon architecture

A main listener thread accepts TCP connections and distributes them round-robin across worker threads (configured with -t, default 4). Each worker runs its own libevent event loop. Connections are non-blocking and sticky to their assigned worker for their lifetime. There is a hard connection limit (configured with -c, default 1024); when it is reached, the listener is temporarily disabled and listen_disabled_num increments. With the default maxconns_fast, memcached returns ERROR Too many open connections and closes each surplus socket; with maxconns_fast disabled, connections instead queue in the OS backlog.

The daemon does not persist data, replicate, or cluster. Treating it as a persistent store, a source of truth, or a strongly consistent system is the most common source of production incidents. Every monitoring signal is downstream of these facts:

  • No persistence means restarts are data-loss events.
  • Clients shard means one hot node indicates uneven key distribution, not a daemon bug.
  • Fixed memory ceiling means evictions are the primary pressure signal, not a defect.

How it works

The slab allocator

The daemon enforces a fixed memory ceiling (the -m flag, default 64 MB). The allocator divides that ceiling into pages of 1 MB. Pages are assigned to slab classes. Each slab class stores items within a specific chunk size range: class 1 handles items up to 96 bytes, class 2 up to 120 bytes, class 3 up to 152 bytes, growing by a configurable factor (the -f flag, default 1.25).

Once a page is assigned to a slab class, it was traditionally never returned. This is the root cause of slab calcification: if your workload’s item size distribution shifts, memory remains locked in classes serving the old distribution while new items evict aggressively in their undersized classes.

Since 1.4.11, slab_reassign allows moving pages between classes (default-on since 1.5.0). Since 1.4.11, slab_automove automates this (mode 1 is the default since 1.5.0). These mitigations reduce calcification but do not eliminate it. The automover is conservative: mode 1 evaluates a default 10-second activity window and then moves at most about one page per second, choosing from an old, idle class for an actively evicting class.

The segmented LRU

Each slab class can maintain a segmented LRU with up to four tiers. HOT, WARM, and COLD are the tiers in the default segmented LRU since 1.5.0:

  • HOT: newly inserted items land here.
  • WARM: items promoted from COLD after a re-access.
  • COLD: eviction candidates. Items that aged out of HOT or WARM without further access.
  • TEMP: items with very short TTLs (opt-in via -o temporary_ttl=<N>) that bypass the full LRU.

A background LRU maintainer thread manages movement between tiers. A separate LRU crawler thread walks the queues reclaiming expired items. Before 1.5.0, a flat per-class LRU was the default and the maintainer was opt-in (-o lru_maintainer).

The segmented design exists to solve a specific problem: scanning workloads (bulk reads of many keys) should not displace the active working set. HOT is probationary; items are never bumped within it. WARM absorbs items that survive COLD. Evictions happen from the tail of COLD. The moves_to_warm and moves_to_cold stats tell you whether the segmentation is working for your workload.

Connections, threads, and the hash table

The main listener thread round-robins connections across worker threads. Each worker has its own libevent loop. There is no key-based routing to workers: the architecture does not create key-to-thread affinity. A hot key accessed through many client connections spreads across workers; a hot key hammered through a single connection concentrates on one worker.

Connection fairness is enforced by the -R limit (max requests per event, default 20), which yields a connection if it sends too many requests in a burst. The conn_yields stat counts these events.

Keys are stored in an expandable hash table. When the table needs to grow, expansion happens in a dedicated background thread with fine-grained locking (since 1.4.x). It is not a stop-the-world operation. The hash_is_expanding stat indicates when this is in progress. During expansion, the old and new tables coexist, so hash_bytes temporarily includes both allocations.

flowchart TD
  mem["-m memory (default 64MB)"] --> pages["1MB pages"]
  pages -->|"assigned to"| cls["Slab classes by item size"]
  cls --> lru["Segmented LRU per class"]
  lru --> hot["HOT - new inserts"]
  lru --> warm["WARM - promoted from COLD on re-access"]
  lru --> cold["COLD - eviction candidates"]
  lru --> temp["TEMP - opt-in short TTL"]
  cold -->|"class full"| evict["Eviction"]

Where it shows up in production

The design creates specific failure archetypes.

Slab imbalance. One slab class is full and evicting while others have free space. Global memory looks fine, but cache effectiveness collapses for items of the saturated size class. This is the most underdiagnosed memcached problem, and the one that global metrics hide best.

Cache stampede. A popular key expires or is evicted. Thousands of requests miss simultaneously and hammer the backend. Memcached itself is fine; the backend is the victim. Low evictions, high misses, and a correlated backend load spike are the signature.

Eviction storm. The working set exceeds cache size. Constant eviction of useful data. Distinguished from healthy turnover by evicted_time: if recently-accessed items are being evicted (low evicted_time in an actively evicting class), the cache is thrashing rather than turning over cold items.

Connection exhaustion. The -c limit is reached. The listener is disabled (listen_disabled_num increments, accepting_conns flips to 0); under the default maxconns_fast, memcached immediately closes surplus sockets. Clients perceive memcached as down or slow. The process itself is healthy and may be idle on existing connections.

Silent degradation. The process is alive (the TCP port accepts connections) but not processing commands. This can happen during deadlock, OOM killer activity, or swap thrashing. A simple port check passes. A command probe (version or stats) fails.

Deployment variants change the monitoring posture:

  • Standalone: single instance, all monitoring is local.
  • Client-side sharded cluster: memcached itself does not cluster. Client libraries shard keys via consistent hashing. One hot node indicates uneven key distribution. Node loss reshuffles the entire keyspace.
  • SASL authenticated: auth failures appear as connection failures, not command failures. SASL uses the binary protocol; separate ASCII authentication is available through -Y. The binary protocol has no formal deprecation status in recent releases.
  • extstore enabled (1.5.4+): external storage for large items. Adds disk I/O as a monitored resource. Strictly opt-in.
  • TLS enabled (1.5.13+): adds TLS handshake overhead and failure stats. Compile-time opt-in.

Tradeoffs and when to use it

Memcached is the right choice when you need a fast, simple, ephemeral cache and your application can tolerate cold-start misses. It is the wrong choice when you need persistence, replication, clustering, or strong consistency.

Common misuses:

  • Treating it as a persistent store. Any data you cannot afford to lose must live elsewhere. Memcached will lose everything on restart, OOM kill, or flush_all.
  • Relying on it for correctness. Atomic counters (incr/decr) wrap silently on overflow and clamp to zero on underflow. Neither produces an error. CAS failures (cas_badval) are silent. If correctness depends on these semantics, you need a different system.
  • Assuming it survives process restarts. Every restart is a cold-cache event. Have a warming strategy and monitor hit ratio recovery.
  • Ignoring client library behavior. Different client libraries handle failures, timeouts, and retries differently. Some silently fall back to a different server, some throw errors, some queue retries. The client library is half the system.

The no-clustering choice is a feature, not a gap. Clustering logic lives in client libraries (consistent hashing rings), which means cluster behavior depends on which client you use and how it is configured. When a node is added or removed, a fraction of keys remaps, producing a temporary burst of misses proportional to the fraction of keys moved.

Signals to watch in production

SignalWhy it mattersWarning sign
Hit ratio (get_hits / (get_hits + get_misses))Primary effectiveness metric. Decline means more traffic hits the backend.Sustained drop of more than 15 percentage points from rolling average.
Eviction rate (evictions)Memory is full in at least one slab class and the LRU is discarding items.Sustained non-zero rate per slab class, not just globally.
Evicted item age (evicted_time per slab class)Discriminates healthy turnover from harmful thrash. Low value means recently-accessed items are being evicted.evicted_time below 300 seconds in any class with active evictions.
Slab class utilization (stats slabs, stats items)Reveals imbalance hiding behind healthy global memory metrics.One class at 100% used_chunks with evictions while others have significant free_chunks.
Connection saturation (curr_connections, accepting_conns, listen_disabled_num)Approaching the hard limit means imminent rejection.curr_connections above 80% of -c, or accepting_conns = 0, or listen_disabled_num increasing.
UptimeReset means restart means total data loss.Unexpected discontinuity in the uptime counter.
cmd_flushEach increment invalidates the entire cache.Any increment in production.
direct_reclaimsThe LRU maintainer is falling behind. Worker threads are evicting inline.Any sustained non-zero rate.

How Netdata helps

  • Per-second visibility into hit ratio, eviction rate, and memory utilization reveals the slab imbalance pattern (one class evicting while others sit idle) before it cascades into backend overload.
  • Correlating memcached signals with backend database metrics on the same timeline turns the eviction-cascade narrative into a single visible story rather than two disconnected alerts firing minutes apart.
  • Anomaly detection on cmd_get, cmd_set, and bytes_written rates catches workload shifts such as item size inflation or traffic spikes that precede eviction storms.
  • Tracking uptime alongside hit ratio recovery curves distinguishes expected post-restart cold-cache behavior from genuine degradation, preventing false pages during planned maintenance.
  • Connection signals (curr_connections, accepting_conns, listen_disabled_num) correlated with client-side metrics isolate connection-pool leaks from real cache pressure.