The only agent that thinks for itself

Autonomous Monitoring with self-learning AI built-in, operating independently across your entire stack.

Unlimited Metrics & Logs
Machine learning & MCP
5% CPU, 150MB RAM
3GB disk, >1 year retention
800+ integrations, zero config
Dashboards, alerts out of the box
> Discover Netdata Agents

Centralized metrics streaming and storage

Aggregate metrics from multiple agents into centralized Parent nodes for unified monitoring across your infrastructure.

Stream from unlimited agents
Long-term data retention
High availability clustering
Data replication & backup
Scalable architecture
Enterprise-grade security
> Learn about Parents

Fully managed cloud platform

Access your monitoring data from anywhere with our SaaS platform. No infrastructure to manage, automatic updates, and global availability.

Zero infrastructure management
99.9% uptime SLA
Global data centers
Automatic updates & patches
Enterprise SSO & RBAC
SOC2 & ISO certified
> Explore Netdata Cloud

Deploy Netdata Cloud in your infrastructure

Run the full Netdata Cloud platform on-premises for complete data sovereignty and compliance with your security policies.

Complete data sovereignty
Air-gapped deployment
Custom compliance controls
Private network integration
Dedicated support team
Kubernetes & Docker support
> Learn about Cloud On-Premises

Powerful, intuitive monitoring interface

Modern, responsive UI built for real-time troubleshooting with customizable dashboards and advanced visualization capabilities.

Real-time chart updates
Customizable dashboards
Dark & light themes
Advanced filtering & search
Responsive on all devices
Collaboration features
> Explore Netdata UI

Monitor on the go

Native iOS and Android apps bring full monitoring capabilities to your mobile device with real-time alerts and notifications.

iOS & Android apps
Push notifications
Touch-optimized interface
Offline data access
Biometric authentication
Widget support
> Download apps

The future of infrastructure observability

See our strategic direction across AI-native observability, full-stack signals, operational intelligence, and enterprise platform maturity.

AI-native observability
Full-stack signal coverage
Operational intelligence
Enterprise platform maturity
Agent releases every 6 weeks
Cloud continuous delivery
> Explore Product Roadmap

Best energy efficiency

True real-time per-second

100% automated zero config

Centralized observability

Multi-year retention

High availability built-in

Zero maintenance

Always up-to-date

Enterprise security

Complete data control

Air-gap ready

Compliance certified

Millisecond responsiveness

Infinite zoom & pan

Works on any device

Native performance

Instant alerts

Monitor anywhere

AI-native observability

Continuous delivery

Open source foundation

80% Faster Incident Resolution

AI-powered troubleshooting from detection, to root cause and blast radius identification, to reporting.

True Real-Time and Simple, even at Scale

Linearly and infinitely scalable full-stack observability, that can be deployed even mid-crisis.

90% Cost Reduction, Full Fidelity

Instead of centralizing the data, Netdata distributes the code, eliminating pipelines and complexity.

See and Map Your Entire Network

Live topology, flow analytics, and SNMP device and trap monitoring — unified with your full-stack observability.

Control Without Surrender

SOC 2 Type 2 certified with every metric kept on your infrastructure.

Integrations

800+ collectors and notification channels, auto-discovered and ready out of the box.

800+ data collectors
Auto-discovery & zero config
Cloud, infra, app protocols
Notifications out of the box
> Explore integrations
Real Results
46% Cost Reduction

Reduced monitoring costs by 46% while cutting staff overhead by 67%.

— Leonardo Antunez, Codyas

Zero Pipeline

No data shipping. No central storage costs. Query at the edge.

From Our Users
"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

No Query Language

Point-and-click troubleshooting. No PromQL, no LogQL, no learning curve.

Enterprise Ready
67% Less Staff, 46% Cost Cut

Enterprise efficiency without enterprise complexity—real ROI from day one.

— Leonardo Antunez, Codyas

SOC 2 Type 2 Certified

Zero data egress. Only metadata reaches the cloud. Your metrics stay on your infrastructure.

Full Coverage
800+ Collectors

Auto-discovered and configured. No manual setup required.

Any Notification Channel

Slack, PagerDuty, Teams, email, webhooks—all built-in.

Built for the People Who Get Paged

Because 3am alerts deserve instant answers, not hour-long hunts.

Every Industry Has Rules. We Master Them.

See how healthcare, finance, and government teams cut monitoring costs 90% while staying audit-ready.

Monitor Any Technology. Configure Nothing.

Install the agent. It already knows your stack.
From Our Users
"A Rare Unicorn"

Netdata gives more than you invest in it. A rare unicorn that obeys the Pareto rule.

— Eduard Porquet Mateu, TMB Barcelona

99% Downtime Reduction

Reduced website downtime by 99% and cloud bill by 30% using Netdata alerts.

— Falkland Islands Government

Real Savings
30% Cloud Cost Reduction

Optimized resource allocation based on Netdata alerts cut cloud spending by 30%.

— Falkland Islands Government

46% Cost Cut

Reduced monitoring staff by 67% while cutting operational costs by 46%.

— Codyas

Real Coverage
"Plugin for Everything"

Netdata has agent capacity or a plugin for everything, including Windows and Kubernetes.

— Eduard Porquet Mateu, TMB Barcelona

"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

Real Speed
Troubleshooting in 30 Seconds

From 2-3 minutes to 30 seconds—instant visibility into any node issue.

— Matthew Artist, Nodecraft

20% Downtime Reduction

20% less downtime and 40% budget optimization from out-of-the-box monitoring.

— Simon Beginn, LANCOM Systems

Pay per Node. Unlimited Everything Else.

One price per node. Unlimited metrics, logs, users, and retention. No per-GB surprises.

Free tier—forever
No metric limits or caps
Retention you control
Cancel anytime
> See pricing plans

What's Your Monitoring Really Costing You?

Most teams overpay by 40-60%. Let's find out why.

Expose hidden metric charges
Calculate tool consolidation
Customers report 30-67% savings
Results in under 60 seconds
> See what you're really paying

Your Infrastructure Is Unique. Let's Talk.

Because monitoring 10 nodes is different from monitoring 10,000.

On-prem & air-gapped deployment
Volume pricing & agreements
Architecture review for your scale
Compliance & security support
> Start a conversation

Monitoring That Sells Itself

Deploy in minutes. Impress clients in hours. Earn recurring revenue for years.

30-second live demos close deals
Zero config = zero support burden
Competitive margins & deal protection
Response in 48 hours
> Apply to partner

Per-Second Metrics at Homelab Prices

Same engine, same dashboards, same ML. Just priced for tinkerers.

Community: Free forever · 5 nodes · non-commercial
Homelab: $90/yr · unlimited nodes · fair usage
> Get the Homelab Plan

$1,000 Per Referral. Unlimited Referrals.

Your colleagues get 10% off. You get 10% commission. Everyone wins.

10% of subscriptions, up to $1,000 each
Track earnings inside Netdata Cloud
PayPal/Venmo payouts in 3-4 weeks
No caps, no complexity
> Get your referral link
Cost Proof
40% Budget Optimization

"Netdata's significant positive impact" — LANCOM Systems

Calculate Your Savings

Compare vs Datadog, Grafana, Dynatrace

Savings Proof
46% Cost Reduction

"Cut costs by 46%, staff by 67%" — Codyas

30% Cloud Bill Savings

"Reduced cloud bill by 30%" — Falkland Islands Gov

Enterprise Proof
"Better Than Combined Alternatives"

"Better observability with Netdata than combining other tools." — TMB Barcelona

Real Engineers, <24h Response

DPA, SLAs, on-prem, volume pricing

Why Partners Win
Demo Live Infrastructure

One command, 30 seconds, real data—no sandbox needed

Zero Tickets, High Margins

Auto-config + per-node pricing = predictable profit

Homelab Ready
Free Video Course

8-episode Netdata tutorial by LearnLinux.tv

76k+ GitHub Stars

3rd most starred monitoring project

Worth Recommending
Product That Delivers

Customers report 40-67% cost cuts, 99% downtime reduction

Zero Risk to Your Rep

Free tier lets them try before they buy

AI Support Assistant, Available 24/7

Nedi has access to all official documentation, source code, and resources. Ask any question about Netdata—responds in your language.

Deployment & configuration
Troubleshooting & sizing
Alerts & notifications
Evidence-based answers
> Ask Nedi now

Never Fight Fires Alone

Docs, community, and expert help—pick your path to resolution.

Learn.netdata.cloud docs
Discord, Forums, GitHub
Premium support available
> Get answers now

60 Seconds to First Dashboard

One command to install. Zero config. 850+ integrations documented.

Linux, Windows, K8s, Docker
Auto-discovers your stack
> Read our documentation

76,000+ Engineers Strong

615+ contributors. 1.5M daily downloads. One mission: simplify observability.

Per-Second. 90% Cheaper. Data Stays Home.

Side-by-side comparisons: costs, real-time granularity, and data sovereignty for every major tool.

See why teams switch from Datadog, Prometheus, Grafana, and more.

> Browse all comparisons
Edge-Native Observability, Born Open Source
Per-second visibility, ML on every metric, and data that never leaves your infrastructure.
Founded in 2016
615+ contributors worldwide
Remote-first, engineering-driven
Open source first
> Read our story
Promises We Publish—and Prove
12 principles backed by open code, independent validation, and measurable outcomes.
Open source, peer-reviewed
Zero config, instant value
Data sovereignty by design
Aligned pricing, no surprises
> See all 12 principles
Edge-Native, AI-Ready, 100% Open
76k+ stars. Full ML, AI, and automation—GPLv3+, not premium add-ons.
76,000+ GitHub stars
GPLv3+ licensed forever
ML on every metric, included
Zero vendor lock-in
> Explore our open source
Build Real-Time Observability for the World
Remote-first team shipping per-second monitoring with ML on every metric.
Remote-first, fully distributed
Open source (76k+ stars)
Challenging technical problems
Your code on millions of systems
> See open roles
Meet the Team Behind Netdata
Conferences, meetups, and tradeshows where you can see Netdata in action and talk to the engineers who build it.
Live demos and deep dives
Book 1-on-1 meetings
Talks and panel sessions
Event recaps and photos
> See all events
Talk to a Netdata Human in <24 Hours
Sales, partnerships, press, or professional services—real engineers, fast answers.
Discuss your observability needs
Pricing and volume discounts
Partnership opportunities
Media and press inquiries
> Book a conversation
Your Data. Your Rules.
On-prem data, cloud control plane, transparent terms.
Trust & Scale
76,000+ GitHub Stars

One of the most popular open-source monitoring projects

SOC 2 Type 2 Certified

Enterprise-grade security and compliance

Data Sovereignty

Your metrics stay on your infrastructure

Validated
University of Amsterdam

"Most energy-efficient monitoring solution" — ICSOC 2023, peer-reviewed

ADASTEC (Autonomous Driving)

"Doesn't miss alerts—mission-critical trust for safety software"

Community Stats
615+ Contributors

Global community improving monitoring for everyone

1.5M+ Downloads/Day

Trusted by teams worldwide

GPLv3+ Licensed

Free forever, fully open source agent

Why Join?
Remote-First

Work from anywhere, async-friendly culture

Impact at Scale

Your work helps millions of systems

$ guides / memcached / memcached-lru-segments-tuning ▌

Operations Guides

Memcached segmented LRU: HOT/WARM/COLD tiers and what the move counters tell you

Since memcached 1.5.0, every slab class runs a segmented LRU instead of a single flat queue. Items land in HOT, WARM, COLD, and optionally TEMP tiers, and a background thread moves them between tiers based on access patterns. The goal is to protect frequently-accessed items from being evicted by one-off scans that touch cold data.

The move counters under stats items describe how well that sorting is working. moves_to_cold counts items aging out of active use, moves_to_warm counts items rescued from COLD by a re-access, and moves_within_lru counts re-ranking within WARM. The ratios between these counters explain why a slab class evicts the way it does, even when global memory looks fine.

Most teams never touch the tier configuration and should not. But if you are debugging per-slab eviction patterns, slab calcification, or a cache with headroom that still thrashes, the move counters are the diagnostic layer beneath the eviction rate.

What it is and why it matters

Before segmented LRU, each slab class had one doubly-linked list. Every access bumped the item to the head; eviction happened at the tail. The problem: a scan that reads many keys once each (a batch job, a crawler, a warm-up script) pushes every scanned item toward the head and evicts genuinely active items at the tail. One pass over cold data can wreck the hit ratio for the entire class.

Segmented LRU, introduced as opt-in in 1.4.23 via -o lru_maintainer and made the default in 1.5.0, splits each slab class into sub-LRUs. A background thread (the LRU maintainer) enforces tier transitions asynchronously, off the request path. A scan that touches cold items once each now drives them toward COLD, not toward the head of a single queue, so they can be evicted without displacing WARM items in the active working set.

How it works

Each slab class maintains up to four sub-LRUs. The LRU maintainer thread iterates them, enforces size and age limits, moves items between tiers, reclaims expired items at tails, and processes asynchronous promotion requests from COLD.

flowchart TD
    NEW["new SET"] --> HOT["HOT: probationary, FIFO"]
    HOT -->|"tail item active"| WARM["WARM: reused items"]
    HOT -->|"tail item inactive"| COLD["COLD: eviction pool"]
    WARM -->|"tail item active, bumped to head"| WARM
    WARM -->|"tail item inactive"| COLD
    COLD -->|"re-accessed: async bump"| WARM
    COLD -->|"evicted at tail"| EVICT["eviction"]

HOT. New items land here on SET. HOT is a FIFO-style probationary queue: items are not bumped to the head on re-access within HOT. When the maintainer reaches the HOT tail, it checks whether that tail item has been re-accessed. Active tail items move to WARM; inactive tail items move to COLD.

WARM. Items that were re-accessed after entering HOT, or rescued from COLD. WARM items are bumped toward the head on re-access, unlike HOT. Inactive items at the WARM tail sink to COLD. WARM buffers workloads where items are read a few times but not continuously.

COLD. The eviction pool. Evictions happen from the COLD tail. If a COLD item is accessed again before eviction, the maintainer queues it for an asynchronous move back to WARM. Under heavy load that bump queue can overflow and some rescues become probabilistic rather than guaranteed.

TEMP. Opt-in via -o temporary_ttl=<N> at startup or lru temp_ttl <N> at runtime. Items with a TTL at or below N seconds bypass HOT, WARM, and COLD entirely. TEMP items are never bumped, never moved between tiers, and are not evictable; the maintainer reaps them on expiry. TEMP exists to stop short-lived items (rate-limit tokens, dedup keys, request-scoped cache entries) polluting the main tiers.

HOT and WARM are each capped at a fraction of per-class memory; current 1.6.x defaults are 20% for HOT and 40% for WARM. A secondary age-based cap also applies: items move out of HOT or WARM when their tail age exceeds a configured factor of COLD’s tail age. These caps stop HOT or WARM monopolizing memory while COLD starves and evictions accelerate.

The move counters

Under stats items, each slab class reports movement counters. These are cumulative since process start; sample twice and compute deltas to get rates.

# Per-slab move counters and tier occupancy
echo "stats items" | nc -q1 localhost 11211 | grep -E "moves_to_(cold|warm)|moves_within_lru|direct_reclaims|number_(hot|warm|cold|temp)"

Three counters describe tier movement, one describes pressure:

CounterWhat it measuresHealthy pattern
moves_to_coldItems sinking from HOT or WARM to COLDSteady flow. Items aging out of active use is normal turnover.
moves_to_warmItems rescued from COLD by a re-accessNon-zero and proportional to your re-access rate. This is segmentation working.
moves_within_lruItems bumped from tail to head within a tier (effectively WARM, since HOT is FIFO)Present when the active working set is being re-ranked in WARM.
direct_reclaimsWorker threads evicting items directly instead of via the maintainerZero. Any non-zero rate means the maintainer fell behind and workers are doing eviction on the request path.

The ratio that matters most is moves_to_warm / moves_to_cold. This is the rescue rate: what fraction of items that reach COLD get pulled back before eviction.

  • High rescue rate: items in COLD are being re-accessed. Segmentation is doing its job: active items survive in WARM, genuinely cold items age out.
  • Low rescue rate: items reach COLD and are never accessed again before eviction. Normal for write-heavy or scan-heavy workloads where most items are touched once. It only becomes a problem if hit ratio is also declining, meaning items you needed were evicted before their next access.
  • Near-zero moves_to_warm with high evictions: the cache is filling with one-off items. Active items get evicted alongside cold ones because the segmentation cannot distinguish them fast enough. This pattern often accompanies scans over a working set larger than the cache.

direct_reclaims is the pressure signal. Normally the maintainer handles all eviction work in the background. When a slab’s SET rate exceeds the maintainer’s ability to move items to COLD and evict from there, worker threads start evicting directly. Each direct reclaim is synchronous work on a worker thread, adding latency to the SET that triggered it. Sustained non-zero direct_reclaims means pressure is acute and the background thread cannot keep up.

Per-slab tier occupancy

stats items also reports how many items sit in each tier per slab class: number_hot, number_warm, number_cold, and (if TEMP is enabled) number_temp. Alongside these, age_hot, age_warm, and age_cold report the age of the oldest item in each tier.

# Tier occupancy and age per slab class
echo "stats items" | nc -q1 localhost 11211 | grep -E "number_(hot|warm|cold|temp)|age_(hot|warm|cold)"

These numbers tell you whether the tiering matches the workload:

  • Most items in WARM with a healthy moves_to_warm rate: the cache has a stable, active working set. Ideal for read-heavy workloads.
  • Most items in COLD: items enter HOT, get one or zero re-accesses, and sink. Expected for write-heavy workloads, but a problem if hit ratio is low.
  • HOT near its cap with low moves_to_warm: new items are arriving faster than they can be promoted or aged out. The class is under write pressure.
  • age_cold very low relative to the typical interval between accesses to popular keys: items are evicted shortly after their last access. That slab class is too small for the working set.

Where this shows up in production

You will reach for these counters in a few specific situations.

Investigating per-slab eviction patterns. When stats items shows one slab class evicting aggressively while others have free chunks, the move counters tell you whether that class is shedding genuinely cold items or losing active ones. High moves_to_warm with rising evicted_time means healthy turnover. Low moves_to_warm with falling evicted_time means thrashing.

Diagnosing scan-heavy workloads. A batch job or crawler that reads thousands of keys once each drives moves_to_cold up and moves_to_warm down. If that coincides with a hit-ratio drop, the scan is displacing the active working set. Enabling TEMP LRU for short-TTL scan entries, or rethinking the scan pattern, may help.

Understanding why adding memory did not help. If a saturated slab class has low moves_to_warm and high direct_reclaims, more memory gives the maintainer headroom but does not change the access pattern. The items being cached are not being re-read. The fix is application-level: cache fewer one-off items, or accept the eviction rate as the cost of the workload.

Deciding whether to enable TEMP LRU. If your workload sets many items with short TTLs (seconds to a minute) and they churn through HOT and COLD without contributing to hit ratio, TEMP LRU isolates them. Set temporary_ttl to cover your shortest-lived items. Be conservative: TEMP items are not evictable, so a threshold set too high can exhaust memory.

Tuning the tiers

Live tuning is available via the lru command. These change runtime behavior on a live cache; lru mode flat in particular alters eviction immediately, so test on a non-production node first.

# Switch between flat and segmented modes (live, changes eviction behavior immediately)
echo "lru mode flat" | nc -q1 localhost 11211
echo "lru mode segmented" | nc -q1 localhost 11211

# Adjust HOT and WARM percentage caps and age factors
echo "lru tune <hot_pct> <warm_pct> <hot_age_factor> <warm_age_factor>" | nc -q1 localhost 11211

# Enable or adjust TEMP LRU
echo "lru temp_ttl <ttl>" | nc -q1 localhost 11211

lru tune adjusts four parameters: the percentage of per-class memory allotted to HOT, the percentage to WARM, and age factors controlling when items move out of HOT or WARM relative to COLD’s tail age. Current defaults are hot_lru_pct=20, warm_lru_pct=40, hot_max_factor=0.20, and warm_max_factor=2.00; stats settings reports the live values.

Most teams should not tune these. The defaults handle typical read-heavy workloads. Adjust the tier caps only with concrete evidence that segmentation is mismatched to your access pattern, such as consistent low rescue rates with declining hit ratio despite adequate total memory.

Version caveat. The lru tune command was inaccessible in memcached 1.6.40 because the rewritten protocol parser compared a command name without checking its length; commit 5f366ccbbfd040a67d2d649f2ae8b5416a2b7677 fixed the regression in 1.6.41. If you run 1.6.40, tuning commands may silently fail or return errors. Upgrade to 1.6.41 or later before attempting tier adjustments.

TEMP LRU risk. Because TEMP items are not evictable, setting temporary_ttl too high fills memory with items that cannot be reclaimed until they expire. Start with a low threshold (a few seconds) and increase only if you see TEMP items expiring before they were useful.

Signals to watch in production

SignalWhy it mattersWarning sign
moves_to_warm / moves_to_cold ratio (per slab)Rescue rate of active items from COLDRatio trending toward zero while evictions climb: nothing is being rescued.
direct_reclaims rate (per slab)Worker threads bypassing the LRU maintainerAny sustained non-zero rate. Pressure exceeds background eviction capacity.
number_cold vs number_hot + number_warm (per slab)Item distribution across tiersAll in COLD: nothing survives probation. All in HOT: write pressure, no re-access.
age_cold (per slab)How long items sit in COLD before evictionVery low relative to popular keys’ access interval: the class is thrashing.
lru_maintainer_juggles (global)How often the maintainer thread woke upSudden sustained increase indicates workload shift or pressure.

These signals sit beneath the per-slab eviction rate and evicted_time in the diagnostic stack. If hit ratio is stable and evicted_time is healthy, you do not need the move counters at all.

How Netdata helps

  • Per-second collection of move counters exposes ratio shifts before they propagate to hit ratio, which is a lagging indicator.
  • Correlating a drop in moves_to_warm with a rise in evictions and a fall in evicted_time pinpoints the moment active items stop being rescued and start being evicted.
  • direct_reclaims as a per-slab alert catches maintainer-thread overload before it manifests as client-visible SET latency.
  • number_hot, number_warm, and number_cold per slab class, alongside eviction rate, show whether the tiering matches the workload or items are pooling in the wrong tier.
  • Anomaly detection on the moves_to_warm / moves_to_cold ratio flags workload shifts the segmentation cannot keep up with, such as a new batch job or a change in key-access patterns.