The only agent that thinks for itself

Autonomous Monitoring with self-learning AI built-in, operating independently across your entire stack.

Unlimited Metrics & Logs
Machine learning & MCP
5% CPU, 150MB RAM
3GB disk, >1 year retention
800+ integrations, zero config
Dashboards, alerts out of the box
> Discover Netdata Agents

Centralized metrics streaming and storage

Aggregate metrics from multiple agents into centralized Parent nodes for unified monitoring across your infrastructure.

Stream from unlimited agents
Long-term data retention
High availability clustering
Data replication & backup
Scalable architecture
Enterprise-grade security
> Learn about Parents

Fully managed cloud platform

Access your monitoring data from anywhere with our SaaS platform. No infrastructure to manage, automatic updates, and global availability.

Zero infrastructure management
99.9% uptime SLA
Global data centers
Automatic updates & patches
Enterprise SSO & RBAC
SOC2 & ISO certified
> Explore Netdata Cloud

Deploy Netdata Cloud in your infrastructure

Run the full Netdata Cloud platform on-premises for complete data sovereignty and compliance with your security policies.

Complete data sovereignty
Air-gapped deployment
Custom compliance controls
Private network integration
Dedicated support team
Kubernetes & Docker support
> Learn about Cloud On-Premises

Powerful, intuitive monitoring interface

Modern, responsive UI built for real-time troubleshooting with customizable dashboards and advanced visualization capabilities.

Real-time chart updates
Customizable dashboards
Dark & light themes
Advanced filtering & search
Responsive on all devices
Collaboration features
> Explore Netdata UI

Monitor on the go

Native iOS and Android apps bring full monitoring capabilities to your mobile device with real-time alerts and notifications.

iOS & Android apps
Push notifications
Touch-optimized interface
Offline data access
Biometric authentication
Widget support
> Download apps

The future of infrastructure observability

See our strategic direction across AI-native observability, full-stack signals, operational intelligence, and enterprise platform maturity.

AI-native observability
Full-stack signal coverage
Operational intelligence
Enterprise platform maturity
Agent releases every 6 weeks
Cloud continuous delivery
> Explore Product Roadmap

Best energy efficiency

True real-time per-second

100% automated zero config

Centralized observability

Multi-year retention

High availability built-in

Zero maintenance

Always up-to-date

Enterprise security

Complete data control

Air-gap ready

Compliance certified

Millisecond responsiveness

Infinite zoom & pan

Works on any device

Native performance

Instant alerts

Monitor anywhere

AI-native observability

Continuous delivery

Open source foundation

80% Faster Incident Resolution

AI-powered troubleshooting from detection, to root cause and blast radius identification, to reporting.

True Real-Time and Simple, even at Scale

Linearly and infinitely scalable full-stack observability, that can be deployed even mid-crisis.

90% Cost Reduction, Full Fidelity

Instead of centralizing the data, Netdata distributes the code, eliminating pipelines and complexity.

See and Map Your Entire Network

Live topology, flow analytics, and SNMP device and trap monitoring — unified with your full-stack observability.

Control Without Surrender

SOC 2 Type 2 certified with every metric kept on your infrastructure.

Integrations

800+ collectors and notification channels, auto-discovered and ready out of the box.

800+ data collectors
Auto-discovery & zero config
Cloud, infra, app protocols
Notifications out of the box
> Explore integrations
Real Results
46% Cost Reduction

Reduced monitoring costs by 46% while cutting staff overhead by 67%.

— Leonardo Antunez, Codyas

Zero Pipeline

No data shipping. No central storage costs. Query at the edge.

From Our Users
"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

No Query Language

Point-and-click troubleshooting. No PromQL, no LogQL, no learning curve.

Enterprise Ready
67% Less Staff, 46% Cost Cut

Enterprise efficiency without enterprise complexity—real ROI from day one.

— Leonardo Antunez, Codyas

SOC 2 Type 2 Certified

Zero data egress. Only metadata reaches the cloud. Your metrics stay on your infrastructure.

Full Coverage
800+ Collectors

Auto-discovered and configured. No manual setup required.

Any Notification Channel

Slack, PagerDuty, Teams, email, webhooks—all built-in.

Built for the People Who Get Paged

Because 3am alerts deserve instant answers, not hour-long hunts.

Every Industry Has Rules. We Master Them.

See how healthcare, finance, and government teams cut monitoring costs 90% while staying audit-ready.

Monitor Any Technology. Configure Nothing.

Install the agent. It already knows your stack.
From Our Users
"A Rare Unicorn"

Netdata gives more than you invest in it. A rare unicorn that obeys the Pareto rule.

— Eduard Porquet Mateu, TMB Barcelona

99% Downtime Reduction

Reduced website downtime by 99% and cloud bill by 30% using Netdata alerts.

— Falkland Islands Government

Real Savings
30% Cloud Cost Reduction

Optimized resource allocation based on Netdata alerts cut cloud spending by 30%.

— Falkland Islands Government

46% Cost Cut

Reduced monitoring staff by 67% while cutting operational costs by 46%.

— Codyas

Real Coverage
"Plugin for Everything"

Netdata has agent capacity or a plugin for everything, including Windows and Kubernetes.

— Eduard Porquet Mateu, TMB Barcelona

"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

Real Speed
Troubleshooting in 30 Seconds

From 2-3 minutes to 30 seconds—instant visibility into any node issue.

— Matthew Artist, Nodecraft

20% Downtime Reduction

20% less downtime and 40% budget optimization from out-of-the-box monitoring.

— Simon Beginn, LANCOM Systems

Pay per Node. Unlimited Everything Else.

One price per node. Unlimited metrics, logs, users, and retention. No per-GB surprises.

Free tier—forever
No metric limits or caps
Retention you control
Cancel anytime
> See pricing plans

What's Your Monitoring Really Costing You?

Most teams overpay by 40-60%. Let's find out why.

Expose hidden metric charges
Calculate tool consolidation
Customers report 30-67% savings
Results in under 60 seconds
> See what you're really paying

Your Infrastructure Is Unique. Let's Talk.

Because monitoring 10 nodes is different from monitoring 10,000.

On-prem & air-gapped deployment
Volume pricing & agreements
Architecture review for your scale
Compliance & security support
> Start a conversation

Monitoring That Sells Itself

Deploy in minutes. Impress clients in hours. Earn recurring revenue for years.

30-second live demos close deals
Zero config = zero support burden
Competitive margins & deal protection
Response in 48 hours
> Apply to partner

Per-Second Metrics at Homelab Prices

Same engine, same dashboards, same ML. Just priced for tinkerers.

Community: Free forever · 5 nodes · non-commercial
Homelab: $90/yr · unlimited nodes · fair usage
> Get the Homelab Plan

$1,000 Per Referral. Unlimited Referrals.

Your colleagues get 10% off. You get 10% commission. Everyone wins.

10% of subscriptions, up to $1,000 each
Track earnings inside Netdata Cloud
PayPal/Venmo payouts in 3-4 weeks
No caps, no complexity
> Get your referral link
Cost Proof
40% Budget Optimization

"Netdata's significant positive impact" — LANCOM Systems

Calculate Your Savings

Compare vs Datadog, Grafana, Dynatrace

Savings Proof
46% Cost Reduction

"Cut costs by 46%, staff by 67%" — Codyas

30% Cloud Bill Savings

"Reduced cloud bill by 30%" — Falkland Islands Gov

Enterprise Proof
"Better Than Combined Alternatives"

"Better observability with Netdata than combining other tools." — TMB Barcelona

Real Engineers, <24h Response

DPA, SLAs, on-prem, volume pricing

Why Partners Win
Demo Live Infrastructure

One command, 30 seconds, real data—no sandbox needed

Zero Tickets, High Margins

Auto-config + per-node pricing = predictable profit

Homelab Ready
Free Video Course

8-episode Netdata tutorial by LearnLinux.tv

76k+ GitHub Stars

3rd most starred monitoring project

Worth Recommending
Product That Delivers

Customers report 40-67% cost cuts, 99% downtime reduction

Zero Risk to Your Rep

Free tier lets them try before they buy

AI Support Assistant, Available 24/7

Nedi has access to all official documentation, source code, and resources. Ask any question about Netdata—responds in your language.

Deployment & configuration
Troubleshooting & sizing
Alerts & notifications
Evidence-based answers
> Ask Nedi now

Never Fight Fires Alone

Docs, community, and expert help—pick your path to resolution.

Learn.netdata.cloud docs
Discord, Forums, GitHub
Premium support available
> Get answers now

60 Seconds to First Dashboard

One command to install. Zero config. 850+ integrations documented.

Linux, Windows, K8s, Docker
Auto-discovers your stack
> Read our documentation

76,000+ Engineers Strong

615+ contributors. 1.5M daily downloads. One mission: simplify observability.

Per-Second. 90% Cheaper. Data Stays Home.

Side-by-side comparisons: costs, real-time granularity, and data sovereignty for every major tool.

See why teams switch from Datadog, Prometheus, Grafana, and more.

> Browse all comparisons
Edge-Native Observability, Born Open Source
Per-second visibility, ML on every metric, and data that never leaves your infrastructure.
Founded in 2016
615+ contributors worldwide
Remote-first, engineering-driven
Open source first
> Read our story
Promises We Publish—and Prove
12 principles backed by open code, independent validation, and measurable outcomes.
Open source, peer-reviewed
Zero config, instant value
Data sovereignty by design
Aligned pricing, no surprises
> See all 12 principles
Edge-Native, AI-Ready, 100% Open
76k+ stars. Full ML, AI, and automation—GPLv3+, not premium add-ons.
76,000+ GitHub stars
GPLv3+ licensed forever
ML on every metric, included
Zero vendor lock-in
> Explore our open source
Build Real-Time Observability for the World
Remote-first team shipping per-second monitoring with ML on every metric.
Remote-first, fully distributed
Open source (76k+ stars)
Challenging technical problems
Your code on millions of systems
> See open roles
Meet the Team Behind Netdata
Conferences, meetups, and tradeshows where you can see Netdata in action and talk to the engineers who build it.
Live demos and deep dives
Book 1-on-1 meetings
Talks and panel sessions
Event recaps and photos
> See all events
Talk to a Netdata Human in <24 Hours
Sales, partnerships, press, or professional services—real engineers, fast answers.
Discuss your observability needs
Pricing and volume discounts
Partnership opportunities
Media and press inquiries
> Book a conversation
Your Data. Your Rules.
On-prem data, cloud control plane, transparent terms.
Trust & Scale
76,000+ GitHub Stars

One of the most popular open-source monitoring projects

SOC 2 Type 2 Certified

Enterprise-grade security and compliance

Data Sovereignty

Your metrics stay on your infrastructure

Validated
University of Amsterdam

"Most energy-efficient monitoring solution" — ICSOC 2023, peer-reviewed

ADASTEC (Autonomous Driving)

"Doesn't miss alerts—mission-critical trust for safety software"

Community Stats
615+ Contributors

Global community improving monitoring for everyone

1.5M+ Downloads/Day

Trusted by teams worldwide

GPLv3+ Licensed

Free forever, fully open source agent

Why Join?
Remote-First

Work from anywhere, async-friendly culture

Impact at Scale

Your work helps millions of systems

$ guides / memcached / memcached-monitoring-maturity-model ▌

Operations Guides

Memcached monitoring maturity model: from survival to expert

Memcached exposes most of its internal state through plain text counters. Most teams stop at hit ratio, global memory, and eviction rate, then discover during an incident that the signal they needed was buried in stats items or /proc/<pid>/status. The four-level model below stages the path from “is the process alive” to “is the segmented LRU moving items between HOT and COLD the way the workload expects”.

Use the levels as a checklist, not a sequence. Each level answers a class of question the previous one could not. Skipping straight to Level 4 without Level 2 fundamentals produces dashboards full of internal counters with no baseline for client impact. The most common gap is between Level 2 and Level 3: operational coverage looks healthy on global metrics while one slab class is silently evicting recently-used items, because the slab allocator partitions memory by item size and the global bytes counter averages across all classes.

The boundaries are deliberate. Survival keeps the service nominally up. Operational proves the cache is doing useful work without quietly swapping or wasting memory. Mature is where slab calcification, LRU maintainer starvation, and write contention become visible. Expert covers deployments running extstore, TLS, very high per-thread throughput, or sharded clusters where cross-node balance matters.

flowchart TD
  L4["Level 4 Expert
per-thread CPU, extstore, TLS, LRU segment dynamics"] L3["Level 3 Mature
per-slab eviction age, direct_reclaims, crawler, automove"] L2["Level 2 Operational
command mix, network bytes, swap, client latency"] L1["Level 1 Survival
process alive, hit ratio, evictions, connections"] L1 --> L2 --> L3 --> L4

Level 1: survival

The minimum signals that tell you memcached is alive and not actively destroying cache effectiveness. If any of these break, you have an incident regardless of what else is monitored.

SignalWhat it answersFirst thing to check
Process livenessIs the daemon responding to commands?echo "version" | nc -q1 localhost 11211
uptimeDid the process restart?STAT uptime from stats; a reset implies full data loss
Hit ratioIs the cache offloading the backend?get_hits / (get_hits + get_misses) from deltas
Eviction rateIs memory under pressure?rate(evictions) from stats
bytes / limit_maxbytesIs global memory near the ceiling?Both fields from stats
curr_connections / max_connectionsAre we approaching the connection limit?Both fields from stats
rejected_connections, listen_disabled_numHave clients been turned away?Counters from stats

The -q1 flag on the liveness probe is openbsd-netcat specific. On systems shipping traditional netcat, substitute an equivalent timeout or use memcapable / a language-specific client. Always probe with an actual command, not a TCP port check: a process that accepts connections but does not respond to commands is effectively down and will pass a naive port check.

Survival-level alerts are blunt by design. A failed version probe sustained for 30 to 60 seconds is a page. A curr_connections ratio above 0.95 or a listen_disabled_num increment is a ticket. Hit ratio and eviction rate are workload-relative and need baselines, but a sudden hit ratio collapse alongside a low uptime is the canonical cold-cache thundering herd pattern and should page only if the backend is also showing load.

One Level 1 trap: global memory utilization can sit at 50% while one slab class is evicting aggressively. That is precisely the gap Level 3 closes.

Level 2: operational

Level 1 told you the service is up. Level 2 tells you whether the cache is doing useful work and whether the process is healthy from the OS perspective.

SignalWhy it matters
cmd_get, cmd_set, cmd_touch, cmd_flush ratesEstablishes the workload profile. Any cmd_flush increment in production is a cache-wide data loss event.
bytes_read, bytes_written ratesNetwork I/O volume. bytes_written approaching NIC capacity is the hidden bottleneck for large-value workloads.
conn_yields rateA connection hit the -R per-event request limit and was forced to yield. Sustained high rates indicate one or two clients dominating the server.
evicted_unfetched, expired_unfetched ratesItems stored but never read before being evicted or expiring. A high evicted_unfetched / evictions ratio means the application is caching write-only data.
reclaimed rateItems whose expired slots were reused. High reclaim with low eviction is the ideal state and confirms the LRU crawler is doing its job.
Process RSS from /proc/<pid>/statusTotal resident memory including slab memory, hash table, connection buffers. RSS more than roughly 1.4x limit_maxbytes suggests overhead growth.
VmSwap from /proc/<pid>/statusMust be zero. Any nonzero value is a production incident; the in-memory cache is now serving some items at disk speed.
Client-observed latencyMemcached exposes no latency histograms. p99 must be measured at the client. p99 above 5ms for same-datacenter traffic is abnormal.
cmd_flush change detection and get_flushedAlert on any cmd_flush increment; get_flushed counts GETs that found flushed items.

The composite connection-exhaustion pattern is the only connection-state condition worth paging on in isolation. It combines accepting_conns = 0, curr_connections / max_connections > 0.98, a positive rate on rejected_connections or time_in_listen_disabled_us, sustained for 5 minutes, with max_connections > 50 to exclude tiny instances. The sustained window filters out deploy-time reconnection storms and brief batch spikes.

Differentiate miss causes. High misses with high evictions means memory pressure. High misses with zero evictions means the application is requesting keys that were never set, or the cache is cold. These have completely different fixes and the global hit ratio counter does not distinguish them.

Level 3: mature

Level 3 is where per-slab instrumentation enters the picture. This is the level most teams never reach, and it is where the most common production incidents become diagnosable. Without it you cannot tell the difference between “the cache is too small” and “one slab class is starved while others have free pages”.

SignalSourceWhat it tells you
Per-slab evicted, age, free_chunks, used_chunks, mem_requestedstats items, stats slabsIdentifies slab calcification: one class at zero free chunks and high evictions while others sit idle.
evicted_time per slab classstats itemsSeconds since last access of the most recently evicted item. Low values (under 300s) with active evictions mean recently-used data is being discarded.
Slab calcification detectionstats slabs cross-class comparisonThe composite pattern: global bytes at 50 to 70% of limit, evictions climbing, one class at 100% used_chunks, others with significant free_chunks.
direct_reclaims per slabstats itemsWorker threads evicting directly instead of letting the LRU maintainer handle it. Sustained non-zero means the maintainer is falling behind.
hash_is_expanding, hash_power_level, hash_bytesstatsHash table growth state. Expansion runs in a background thread but doubles hash table memory temporarily.
crawler_reclaimed, crawler_items_checked ratiostatsLRU crawler effectiveness. A dropping reclaim rate alongside active evictions suggests the crawler is disabled or stuck.
Slab automove status and activitystats settings, stats slabsWhether slab_automove is on (default mode 1 since 1.5.0). Mode 2 is aggressive and not recommended for long-term use.
cas_badval ratestatsCAS validation failures. cas_badval / (cas_hits + cas_misses + cas_badval) above 10% indicates write contention on hot keys.
incr_misses, decr_misses ratesstatsCounter keys being evicted or expired between increments. Breaks rate limiters and distributed counters.
response_obj_oomstats (1.6.0+)Response buffer allocation failures forcing connection closes. Separate from item eviction memory; can occur with zero evictions.
store_too_large, store_no_memorystatsstore_too_large indicates oversized items (application serialization bug). store_no_memory indicates -M no-eviction mode is refusing writes.
auth_cmds, auth_errors (if auth enabled)statsSASL or ASCII auth activity. SASL requires the deprecated binary protocol; zero errors with auth disabled does not mean access is controlled.

The single most important Level 3 skill is reading evicted_time correctly. The counter is per-slab, only meaningful when that slab is actively evicting, and reports only the age of the last evicted item, not an average. A 3-day-old evicted item is healthy cache turnover. A 10-second-old evicted item means the cache is actively thrashing and the items being discarded would have been hit. Distinguishing the two is impossible from the global evictions counter alone.

Level 3 is also where flush_all becomes diagnosable rather than mysterious. cmd_flush increments, and the post-flush state shows curr_items declining and hit ratio dropping. The flush itself does not lock the server and does not free memory immediately; items are lazily invalidated on next access.

Level 4: expert

Level 4 is for deployments where the segmented LRU’s internal dynamics, per-thread CPU saturation, extstore, TLS, or cross-instance balance matter. These signals are added after a specific class of incident has already happened once.

SignalWhat it reveals
moves_to_cold, moves_to_warm, moves_within_lru ratesWhether the segmented LRU is moving items between HOT, WARM, COLD, and TEMP the way the workload expects. moves_to_warm / moves_to_cold indicates the rescue rate for re-accessed items.
Per-slab fragmentation: mem_requested / (used_chunks * chunk_size)Internal slab waste. Below 0.5 for a class means items are using less than half their allocated chunk, suggesting the -f growth factor needs tuning.
Cross-instance balance in sharded clustersPer-node command rates, hit ratios, and evictions. One hot node indicates uneven key distribution; client libraries do the sharding, so memcached itself has no cluster-wide view.
Per-thread CPU utilizationSingle worker thread at 100% while others sit idle. Aggregate rusage averages across threads and hides this. Requires OS-level per-thread monitoring of the memcached process.
TCP connection state via ss -tnTIME_WAIT and CLOSE_WAIT accumulation indicating client-side connection handling problems. Memcached stats do not expose this.
lrutail_reflocked rateItems at the LRU tail with nonzero refcount, blocking eviction. High values indicate large-value reads competing with memory pressure.
Item size distribution via -o track_sizes (available since 1.4.27; runtime enable/disable removed in 1.6.25)Live histogram of item sizes, useful for sizing the slab growth factor and detecting serialization drift.
total_malloced vs limit_maxbytesHow much of the configured memory has actually been claimed by slabs. A low ratio means the working set has not yet demanded the full allocation.

Conditional Level 4 signals apply only to specific deployment variants:

  • extstore enabled (feature introduced in 1.5.4; compiled in by default since 1.6.0 but still enabled at startup): get_extstore, get_aborted_extstore, get_oom_extstore, recache_from_extstore, extstore_page_allocs, extstore_page_evictions, extstore_objects_evicted/read/written/used, extstore_bytes_*, extstore_io_queue. extstore adds disk I/O as a monitored resource.
  • TLS enabled (experimental support introduced in 1.5.13; TLS stats vary by version): ssl_handshake_errors, ssl_proto_errors, ssl_new_sessions, time_since_server_cert_refresh. TLS splits messages above 16KB by default; buffer size is tunable via -o ssl_wbuf_size=. Compile-time opt-in.
  • SASL authenticated: auth_cmds and auth_errors. SASL requires the deprecated binary protocol, and auth_errors = 0 with auth disabled does not mean access is controlled.

Expert level also includes version-aware behavior changes that affect what the stats mean. Segmented LRU is default since 1.5.0. slab_automove mode 1 is default since 1.5.0. UDP is disabled by default since 1.5.6, the release that mitigated CVE-2018-1000115. The binary protocol is officially deprecated since 1.6.0 in favor of the meta and text protocols. Treating all versions identically will produce wrong conclusions about which features are active.

How Netdata helps

  • Per-second deltas on cmd_get, cmd_set, cmd_touch, cmd_flush, and the full hit/miss family mean hit ratio is computed from rate windows rather than lifetime cumulative counters.
  • Per-slab charts surface evicted, age, free_chunks, used_chunks, and mem_requested per class, making slab calcification visible without manually diffing stats slabs output between polls.
  • evicted_time is tracked per slab class alongside eviction rate, so the “evicting cold items vs evicting recently-used items” distinction shows up as a correlated pair.
  • OS-level signals (RSS, VmSwap, per-thread CPU, NIC utilization, TCP state) are collected from the same agent, which matters most for silent-degradation patterns where memcached’s own stats look healthy but the process is partially swapped or a single worker thread is pinned.
  • Anomaly detection on rate-of-change signals catches the gradual hit-ratio decline and the slow per-slab free_chunks trend that absolute thresholds miss.
  • cmd_flush increments and get_flushed rates land on the same timeline, so post-flush cold-cache events are diagnosable from a single view rather than a log scrape plus a manual stats poll.