The only agent that thinks for itself

Autonomous Monitoring with self-learning AI built-in, operating independently across your entire stack.

Unlimited Metrics & Logs
Machine learning & MCP
5% CPU, 150MB RAM
3GB disk, >1 year retention
800+ integrations, zero config
Dashboards, alerts out of the box
> Discover Netdata Agents

Centralized metrics streaming and storage

Aggregate metrics from multiple agents into centralized Parent nodes for unified monitoring across your infrastructure.

Stream from unlimited agents
Long-term data retention
High availability clustering
Data replication & backup
Scalable architecture
Enterprise-grade security
> Learn about Parents

Fully managed cloud platform

Access your monitoring data from anywhere with our SaaS platform. No infrastructure to manage, automatic updates, and global availability.

Zero infrastructure management
99.9% uptime SLA
Global data centers
Automatic updates & patches
Enterprise SSO & RBAC
SOC2 & ISO certified
> Explore Netdata Cloud

Deploy Netdata Cloud in your infrastructure

Run the full Netdata Cloud platform on-premises for complete data sovereignty and compliance with your security policies.

Complete data sovereignty
Air-gapped deployment
Custom compliance controls
Private network integration
Dedicated support team
Kubernetes & Docker support
> Learn about Cloud On-Premises

Powerful, intuitive monitoring interface

Modern, responsive UI built for real-time troubleshooting with customizable dashboards and advanced visualization capabilities.

Real-time chart updates
Customizable dashboards
Dark & light themes
Advanced filtering & search
Responsive on all devices
Collaboration features
> Explore Netdata UI

Monitor on the go

Native iOS and Android apps bring full monitoring capabilities to your mobile device with real-time alerts and notifications.

iOS & Android apps
Push notifications
Touch-optimized interface
Offline data access
Biometric authentication
Widget support
> Download apps

The future of infrastructure observability

See our strategic direction across AI-native observability, full-stack signals, operational intelligence, and enterprise platform maturity.

AI-native observability
Full-stack signal coverage
Operational intelligence
Enterprise platform maturity
Agent releases every 6 weeks
Cloud continuous delivery
> Explore Product Roadmap

Best energy efficiency

True real-time per-second

100% automated zero config

Centralized observability

Multi-year retention

High availability built-in

Zero maintenance

Always up-to-date

Enterprise security

Complete data control

Air-gap ready

Compliance certified

Millisecond responsiveness

Infinite zoom & pan

Works on any device

Native performance

Instant alerts

Monitor anywhere

AI-native observability

Continuous delivery

Open source foundation

80% Faster Incident Resolution

AI-powered troubleshooting from detection, to root cause and blast radius identification, to reporting.

True Real-Time and Simple, even at Scale

Linearly and infinitely scalable full-stack observability, that can be deployed even mid-crisis.

90% Cost Reduction, Full Fidelity

Instead of centralizing the data, Netdata distributes the code, eliminating pipelines and complexity.

See and Map Your Entire Network

Live topology, flow analytics, and SNMP device and trap monitoring — unified with your full-stack observability.

Control Without Surrender

SOC 2 Type 2 certified with every metric kept on your infrastructure.

Integrations

800+ collectors and notification channels, auto-discovered and ready out of the box.

800+ data collectors
Auto-discovery & zero config
Cloud, infra, app protocols
Notifications out of the box
> Explore integrations
Real Results
46% Cost Reduction

Reduced monitoring costs by 46% while cutting staff overhead by 67%.

— Leonardo Antunez, Codyas

Zero Pipeline

No data shipping. No central storage costs. Query at the edge.

From Our Users
"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

No Query Language

Point-and-click troubleshooting. No PromQL, no LogQL, no learning curve.

Enterprise Ready
67% Less Staff, 46% Cost Cut

Enterprise efficiency without enterprise complexity—real ROI from day one.

— Leonardo Antunez, Codyas

SOC 2 Type 2 Certified

Zero data egress. Only metadata reaches the cloud. Your metrics stay on your infrastructure.

Full Coverage
800+ Collectors

Auto-discovered and configured. No manual setup required.

Any Notification Channel

Slack, PagerDuty, Teams, email, webhooks—all built-in.

Built for the People Who Get Paged

Because 3am alerts deserve instant answers, not hour-long hunts.

Every Industry Has Rules. We Master Them.

See how healthcare, finance, and government teams cut monitoring costs 90% while staying audit-ready.

Monitor Any Technology. Configure Nothing.

Install the agent. It already knows your stack.
From Our Users
"A Rare Unicorn"

Netdata gives more than you invest in it. A rare unicorn that obeys the Pareto rule.

— Eduard Porquet Mateu, TMB Barcelona

99% Downtime Reduction

Reduced website downtime by 99% and cloud bill by 30% using Netdata alerts.

— Falkland Islands Government

Real Savings
30% Cloud Cost Reduction

Optimized resource allocation based on Netdata alerts cut cloud spending by 30%.

— Falkland Islands Government

46% Cost Cut

Reduced monitoring staff by 67% while cutting operational costs by 46%.

— Codyas

Real Coverage
"Plugin for Everything"

Netdata has agent capacity or a plugin for everything, including Windows and Kubernetes.

— Eduard Porquet Mateu, TMB Barcelona

"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

Real Speed
Troubleshooting in 30 Seconds

From 2-3 minutes to 30 seconds—instant visibility into any node issue.

— Matthew Artist, Nodecraft

20% Downtime Reduction

20% less downtime and 40% budget optimization from out-of-the-box monitoring.

— Simon Beginn, LANCOM Systems

Pay per Node. Unlimited Everything Else.

One price per node. Unlimited metrics, logs, users, and retention. No per-GB surprises.

Free tier—forever
No metric limits or caps
Retention you control
Cancel anytime
> See pricing plans

What's Your Monitoring Really Costing You?

Most teams overpay by 40-60%. Let's find out why.

Expose hidden metric charges
Calculate tool consolidation
Customers report 30-67% savings
Results in under 60 seconds
> See what you're really paying

Your Infrastructure Is Unique. Let's Talk.

Because monitoring 10 nodes is different from monitoring 10,000.

On-prem & air-gapped deployment
Volume pricing & agreements
Architecture review for your scale
Compliance & security support
> Start a conversation

Monitoring That Sells Itself

Deploy in minutes. Impress clients in hours. Earn recurring revenue for years.

30-second live demos close deals
Zero config = zero support burden
Competitive margins & deal protection
Response in 48 hours
> Apply to partner

Per-Second Metrics at Homelab Prices

Same engine, same dashboards, same ML. Just priced for tinkerers.

Community: Free forever · 5 nodes · non-commercial
Homelab: $90/yr · unlimited nodes · fair usage
> Get the Homelab Plan

$1,000 Per Referral. Unlimited Referrals.

Your colleagues get 10% off. You get 10% commission. Everyone wins.

10% of subscriptions, up to $1,000 each
Track earnings inside Netdata Cloud
PayPal/Venmo payouts in 3-4 weeks
No caps, no complexity
> Get your referral link
Cost Proof
40% Budget Optimization

"Netdata's significant positive impact" — LANCOM Systems

Calculate Your Savings

Compare vs Datadog, Grafana, Dynatrace

Savings Proof
46% Cost Reduction

"Cut costs by 46%, staff by 67%" — Codyas

30% Cloud Bill Savings

"Reduced cloud bill by 30%" — Falkland Islands Gov

Enterprise Proof
"Better Than Combined Alternatives"

"Better observability with Netdata than combining other tools." — TMB Barcelona

Real Engineers, <24h Response

DPA, SLAs, on-prem, volume pricing

Why Partners Win
Demo Live Infrastructure

One command, 30 seconds, real data—no sandbox needed

Zero Tickets, High Margins

Auto-config + per-node pricing = predictable profit

Homelab Ready
Free Video Course

8-episode Netdata tutorial by LearnLinux.tv

76k+ GitHub Stars

3rd most starred monitoring project

Worth Recommending
Product That Delivers

Customers report 40-67% cost cuts, 99% downtime reduction

Zero Risk to Your Rep

Free tier lets them try before they buy

AI Support Assistant, Available 24/7

Nedi has access to all official documentation, source code, and resources. Ask any question about Netdata—responds in your language.

Deployment & configuration
Troubleshooting & sizing
Alerts & notifications
Evidence-based answers
> Ask Nedi now

Never Fight Fires Alone

Docs, community, and expert help—pick your path to resolution.

Learn.netdata.cloud docs
Discord, Forums, GitHub
Premium support available
> Get answers now

60 Seconds to First Dashboard

One command to install. Zero config. 850+ integrations documented.

Linux, Windows, K8s, Docker
Auto-discovers your stack
> Read our documentation

76,000+ Engineers Strong

615+ contributors. 1.5M daily downloads. One mission: simplify observability.

Per-Second. 90% Cheaper. Data Stays Home.

Side-by-side comparisons: costs, real-time granularity, and data sovereignty for every major tool.

See why teams switch from Datadog, Prometheus, Grafana, and more.

> Browse all comparisons
Edge-Native Observability, Born Open Source
Per-second visibility, ML on every metric, and data that never leaves your infrastructure.
Founded in 2016
615+ contributors worldwide
Remote-first, engineering-driven
Open source first
> Read our story
Promises We Publish—and Prove
12 principles backed by open code, independent validation, and measurable outcomes.
Open source, peer-reviewed
Zero config, instant value
Data sovereignty by design
Aligned pricing, no surprises
> See all 12 principles
Edge-Native, AI-Ready, 100% Open
76k+ stars. Full ML, AI, and automation—GPLv3+, not premium add-ons.
76,000+ GitHub stars
GPLv3+ licensed forever
ML on every metric, included
Zero vendor lock-in
> Explore our open source
Build Real-Time Observability for the World
Remote-first team shipping per-second monitoring with ML on every metric.
Remote-first, fully distributed
Open source (76k+ stars)
Challenging technical problems
Your code on millions of systems
> See open roles
Meet the Team Behind Netdata
Conferences, meetups, and tradeshows where you can see Netdata in action and talk to the engineers who build it.
Live demos and deep dives
Book 1-on-1 meetings
Talks and panel sessions
Event recaps and photos
> See all events
Talk to a Netdata Human in <24 Hours
Sales, partnerships, press, or professional services—real engineers, fast answers.
Discuss your observability needs
Pricing and volume discounts
Partnership opportunities
Media and press inquiries
> Book a conversation
Your Data. Your Rules.
On-prem data, cloud control plane, transparent terms.
Trust & Scale
76,000+ GitHub Stars

One of the most popular open-source monitoring projects

SOC 2 Type 2 Certified

Enterprise-grade security and compliance

Data Sovereignty

Your metrics stay on your infrastructure

Validated
University of Amsterdam

"Most energy-efficient monitoring solution" — ICSOC 2023, peer-reviewed

ADASTEC (Autonomous Driving)

"Doesn't miss alerts—mission-critical trust for safety software"

Community Stats
615+ Contributors

Global community improving monitoring for everyone

1.5M+ Downloads/Day

Trusted by teams worldwide

GPLv3+ Licensed

Free forever, fully open source agent

Why Join?
Remote-First

Work from anywhere, async-friendly culture

Impact at Scale

Your work helps millions of systems

$ guides / memcached / memcached-slab-automove ▌

Operations Guides

Memcached slab_automove and slab_reassign: rebalancing pages between classes

Memcached’s slab allocator divides its memory budget into 1MB pages and permanently assigns each page to a slab class based on item size. Once assigned, a page traditionally stayed locked to that class forever. If the workload’s item-size distribution shifted after deployment, you got slab calcification: one class full and evicting while others sat idle with free chunks. Global memory utilization looked healthy while cache effectiveness collapsed for the saturated size range.

Two mechanisms address this. slab_reassign (introduced in 1.4.11) moves individual pages between slab classes at runtime. slab_automove (also 1.4.11) automates that decision with a background thread that watches eviction patterns and relocates pages accordingly. Mode 1 is the conservative mover; mode 2 is aggressive and not recommended for sustained use.

These are not free. Page moves are destructive to the source class. The automover cannot help when every slab class is evicting. And the underlying cause, a mismatch between your item-size distribution and the slab class boundaries, often requires a structural fix via the growth factor (-f) on the next restart. This article covers when to rely on runtime rebalancing, when it cannot help, and when you need the restart-time lever instead.

Both are enabled by default since 1.5.0 (slab_reassign on, slab_automove mode 1) and remain so through current 1.6.x; on 1.4.x both were opt-in. The stats settings output shows the effective state.

What slab_reassign and slab_automove actually do

The slab allocator pre-allocates a fixed memory budget (the -m flag, default 64MB) at startup. It divides this into 1MB pages and assigns each page to a slab class. Each class stores items within a specific chunk-size range. Class 1 handles items up to 96 bytes, class 2 up to 120 bytes, class 3 up to 152 bytes, growing by the configurable factor -f (default 1.25).

Before slab_reassign, a page assigned to a class stayed there permanently. If the application’s serialization format changed, or a new feature started caching differently-sized values, the old slab classes held pages they no longer needed while the new size range starved. The only fix was a restart, which wiped all cached data.

slab_reassign breaks that permanence. It lets you, or the automover, move a 1MB page from one slab class to another at runtime. The move is destructive: every item stored in that page within the source class is evicted. Since 1.4.25, the automover attempts to rescue still-valid items from the source page by relocating them to other pages in the same class before the move completes. Items that cannot be rescued are dropped. Freed memory can also be reclaimed into a global page pool and reassigned to new slab classes as needed.

How the automover decides

The automover runs in a background thread and evaluates slab classes on a sliding window. The window tracks eviction activity per class over time.

Mode 0 (disabled): No automatic page movement. Pages stay where they were initially assigned.

Mode 1 (conservative, the default since 1.5.0): The automover evaluates every second over a rolling window of slab_automove_window samples (default 10; -o slab_automove_window). A class becomes a move target if it is evicting and its evictions filled more than half the window’s samples, or it drew over a quarter of all window evictions. The source class is the one with the oldest average item age holding more than 2 pages (MIN_PAGES_FOR_SOURCE = 2 in the source). A page moves when the target’s item age is below slab_automove_ratio (default 0.80) of the source’s age. Spare capacity — more than 2.5 pages’ worth of free chunks — is returned to a global page pool instead. This is deliberately slow. Third-party distributions rarely change this default, but stats settings shows the effective mode.

Mode 2 (aggressive): Moves a page on every eviction event. It reacts instantly to pressure but causes significant latency jitter because each move locks items, rescues what it can, and drops the rest. Not recommended for long-term use.

flowchart TD
    A["Automover evaluates every second"] --> B{"Target class: evicting with young item age?"}
    B -->|yes| C{"Source class: oldest average age, > 2 pages?"}
    B -->|no| Z["No action this cycle"]
    C -->|found| D["Move 1 page from source to target"]
    C -->|"all classes evicting"| E["Cannot help: no donor available"]
    D --> F["Rescue valid items to other pages in source class"]
    F --> G["Drop non-rescued items"]
    G --> H["Reassign page to target class"]

The key constraint is that the automover can only take pages from classes that are not themselves under eviction pressure. If every active slab class is evicting, there is no donor, and the automover has nothing to do. This is the most common reason operators enable rebalancing and see no effect.

Enabling and tuning slab rebalancing

Check whether the automover is active:

# Check automove mode and slab_reassign setting
echo "stats settings" | nc localhost 11211 | grep slab

Enable conservative mode at runtime (no restart needed):

# Enable automove mode 1 (conservative)
echo "slabs automove 1" | nc localhost 11211

Disable it:

# Disable automove
echo "slabs automove 0" | nc localhost 11211

For a manual one-shot page move, identify the source and target slab classes from stats slabs, then issue the reassign command directly:

# Identify which classes have free pages and which are saturated
echo "stats slabs" | nc localhost 11211 | grep -E "(total_pages|free_chunks|used_chunks)"
echo "stats items" | nc localhost 11211 | grep -E "(evicted|evicted_time)"

# Move one page from class 5 to class 12 (destructive to class 5 items in that page).
# The command returns a status string; check it before assuming the move happened.
echo "slabs reassign 5 12" | nc localhost 11211

The response is OK when the move is scheduled. Other responses carry a message: BUSY (a move is already in progress), BADCLASS (invalid class id), NOSPARE (source has no spare pages), NOTFULL (destination class must be full), UNSAFE (source cannot move a page right now), or SAME (source and destination identical).

Manual reassignment is useful when the automover is off or when you need to force a specific move the automover would not make on its own. The tradeoff is the same: items in the source page are evicted unless they can be rescued. For targeted relief during an active slab imbalance incident, a manual move can be faster than waiting for the conservative automover to act. Do not run a manual reassign against a class that is itself evicting; you are stealing memory from a saturated range.

When slab rebalancing cannot help

Every slab class is evicting. The automover’s algorithm requires a donor class with zero recent evictions. If all classes are under pressure, there is no page to take. This typically means the working set genuinely exceeds the allocated memory, and the fix is more RAM or fewer cached items, not page shuffling.

The automover anti-favours a specific class. A documented production issue showed the automover persistently starving one slab class (memcached issue #677, observed on 1.5.12 in May 2020 and closed as fixed in September 2020 after the automover was reworked into the window-and-ratio algorithm described above). The automover would donate pages away from that class to serve other classes, then the class would evict heavily. Manual reassignment was ineffective because the automover immediately took the page back. The counters told the story: slab_reassign_busy_items (2.4 billion) and slab_reassign_evictions_nomem (2.4 billion) ran into the billions, indicating the automover was constantly attempting moves that failed. The workaround was to explicitly steal pages from the class with the most pages, or disable the automover entirely and manage moves manually.

Reassignment is slow relative to workload changes. The automover moves one page per 10 seconds. If a workload shift causes sudden concentrated pressure on a class that needs many pages, the automover cannot keep up. A single page move under contention can require a large number of request cycles to complete while the rescuer waits for locked items to drain. For gradual drift, this is fine. For sudden shifts, it is too slow.

Old versions have a worse automover. The automover in 1.4.x is significantly less effective than in 1.5.x and 1.6.x. Operators running older versions who observe that automove “doesn’t seem to do anything” should consider upgrading before concluding the feature is useless. The 1.4.25 release brought major improvements (global page pool, item rescue), and 1.5.0+ refined the algorithm further.

The structural fix: growth factor (-f)

When slab imbalance recurs after every restart, or when the automover cannot keep up with a persistent distribution mismatch, the root cause is usually the chunk-size boundaries determined by the growth factor.

The -f flag (default 1.25) controls how slab class boundaries grow. With -f 1.25, class boundaries are roughly 96, 120, 152, 192 bytes, and so on. Lower values (for example 1.10 or 1.15) produce more classes with finer granularity between sizes, reducing internal fragmentation (items stored in oversized chunks) but creating more classes competing for the same fixed pool of pages. Higher values produce fewer classes, which means each class gets more pages but items waste more space within their chunks.

# Check current growth factor
echo "stats settings" | nc localhost 11211 | grep "STAT growth_factor"

Changing -f requires a restart, which means total data loss. The decision should be based on analysis of the item-size distribution under the current workload. Use stats slabs to identify which classes are perpetually starved and which have excess pages. If the starved classes are adjacent (for example, 120-byte and 152-byte items both evicting because they share a boundary), a lower growth factor creates more classes in that range and spreads the load.

The tradeoff is straightforward: lower -f reduces fragmentation at the cost of more class competition for pages. With slab_automove enabled, the server can redistribute pages over time to match the new class layout, but the initial allocation on restart will reflect the new boundaries.

Signals to watch

SignalWhy it mattersWarning sign
slabs_moved (stats)Total pages moved by automove or manual reassignFlat while one class evicts and another has free pages means automove is off or cannot find a valid donor
slab_reassign_running (stats)Whether the background reassign thread is activePersistently 1 suggests a stuck or extremely slow reassignment
slab_reassign_busy_items (stats)Items that could not be relocated during a move because they were locked by active requestsVery high values indicate the source class has hot items being constantly accessed, making moves mostly destructive
slab_reassign_rescues (stats)Items successfully relocated to other pages before a page moveLow rescues relative to busy_items means most items in moved pages are being dropped
slab_reassign_evictions_nomem (slab_reassign_busy_nomem since 1.6.34)Items evicted because no memory was available for rescueHigh values mean page moves are causing collateral eviction beyond the source page
slab_global_page_pool (stats)Pages in the global reclaim pool available for reassignmentZero means no spare pages for new classes to draw from
Per-slab evicted + free_chunks (stats items, stats slabs)Identifies which classes are starving versus which have spare capacityOne class evicting while another has free_chunks > 0 is the classic imbalance signal
evicted_time per slab (stats items)Age of the most recently evicted item in each classLow evicted_time (under 300 seconds) in a class with active evictions indicates harmful thrashing, not healthy turnover
slab_automove (stats settings)Current automove mode (0, 1, or 2)Mode 0 with active slab imbalance means rebalancing is available but not enabled

How Netdata helps

Netdata surfaces per-slab eviction counts, evicted_time, free_chunks, and used_chunks as per-second time series, which lets you see slab imbalance developing before it becomes an incident.

  • Correlate slabs_moved with per-slab eviction rates to verify the automover is actually addressing the imbalance, or to confirm it is stuck because all classes are evicting.
  • Track slab_reassign_busy_items and slab_reassign_rescues rates to distinguish productive page moves (high rescues) from destructive ones (high busy_items, low rescues).
  • Watch per-slab evicted_time trends to detect when evictions shift from healthy cold-item turnover to harmful thrashing of recently-accessed data.
  • Alert on slab_automove setting changes to catch unexpected disabling of the automover.
  • Correlate automove activity with client-observed latency spikes to identify whether mode 2 aggressive moves are causing jitter.
  • Compare global memory utilization against per-slab utilization to surface the gap that makes slab calcification invisible at the aggregate level.