The only agent that thinks for itself

Autonomous Monitoring with self-learning AI built-in, operating independently across your entire stack.

Unlimited Metrics & Logs
Machine learning & MCP
5% CPU, 150MB RAM
3GB disk, >1 year retention
800+ integrations, zero config
Dashboards, alerts out of the box
> Discover Netdata Agents

Centralized metrics streaming and storage

Aggregate metrics from multiple agents into centralized Parent nodes for unified monitoring across your infrastructure.

Stream from unlimited agents
Long-term data retention
High availability clustering
Data replication & backup
Scalable architecture
Enterprise-grade security
> Learn about Parents

Fully managed cloud platform

Access your monitoring data from anywhere with our SaaS platform. No infrastructure to manage, automatic updates, and global availability.

Zero infrastructure management
99.9% uptime SLA
Global data centers
Automatic updates & patches
Enterprise SSO & RBAC
SOC2 & ISO certified
> Explore Netdata Cloud

Deploy Netdata Cloud in your infrastructure

Run the full Netdata Cloud platform on-premises for complete data sovereignty and compliance with your security policies.

Complete data sovereignty
Air-gapped deployment
Custom compliance controls
Private network integration
Dedicated support team
Kubernetes & Docker support
> Learn about Cloud On-Premises

Powerful, intuitive monitoring interface

Modern, responsive UI built for real-time troubleshooting with customizable dashboards and advanced visualization capabilities.

Real-time chart updates
Customizable dashboards
Dark & light themes
Advanced filtering & search
Responsive on all devices
Collaboration features
> Explore Netdata UI

Monitor on the go

Native iOS and Android apps bring full monitoring capabilities to your mobile device with real-time alerts and notifications.

iOS & Android apps
Push notifications
Touch-optimized interface
Offline data access
Biometric authentication
Widget support
> Download apps

The future of infrastructure observability

See our strategic direction across AI-native observability, full-stack signals, operational intelligence, and enterprise platform maturity.

AI-native observability
Full-stack signal coverage
Operational intelligence
Enterprise platform maturity
Agent releases every 6 weeks
Cloud continuous delivery
> Explore Product Roadmap

Best energy efficiency

True real-time per-second

100% automated zero config

Centralized observability

Multi-year retention

High availability built-in

Zero maintenance

Always up-to-date

Enterprise security

Complete data control

Air-gap ready

Compliance certified

Millisecond responsiveness

Infinite zoom & pan

Works on any device

Native performance

Instant alerts

Monitor anywhere

AI-native observability

Continuous delivery

Open source foundation

80% Faster Incident Resolution

AI-powered troubleshooting from detection, to root cause and blast radius identification, to reporting.

True Real-Time and Simple, even at Scale

Linearly and infinitely scalable full-stack observability, that can be deployed even mid-crisis.

90% Cost Reduction, Full Fidelity

Instead of centralizing the data, Netdata distributes the code, eliminating pipelines and complexity.

See and Map Your Entire Network

Live topology, flow analytics, and SNMP device and trap monitoring — unified with your full-stack observability.

Control Without Surrender

SOC 2 Type 2 certified with every metric kept on your infrastructure.

Integrations

800+ collectors and notification channels, auto-discovered and ready out of the box.

800+ data collectors
Auto-discovery & zero config
Cloud, infra, app protocols
Notifications out of the box
> Explore integrations
Real Results
46% Cost Reduction

Reduced monitoring costs by 46% while cutting staff overhead by 67%.

— Leonardo Antunez, Codyas

Zero Pipeline

No data shipping. No central storage costs. Query at the edge.

From Our Users
"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

No Query Language

Point-and-click troubleshooting. No PromQL, no LogQL, no learning curve.

Enterprise Ready
67% Less Staff, 46% Cost Cut

Enterprise efficiency without enterprise complexity—real ROI from day one.

— Leonardo Antunez, Codyas

SOC 2 Type 2 Certified

Zero data egress. Only metadata reaches the cloud. Your metrics stay on your infrastructure.

Full Coverage
800+ Collectors

Auto-discovered and configured. No manual setup required.

Any Notification Channel

Slack, PagerDuty, Teams, email, webhooks—all built-in.

Built for the People Who Get Paged

Because 3am alerts deserve instant answers, not hour-long hunts.

Every Industry Has Rules. We Master Them.

See how healthcare, finance, and government teams cut monitoring costs 90% while staying audit-ready.

Monitor Any Technology. Configure Nothing.

Install the agent. It already knows your stack.
From Our Users
"A Rare Unicorn"

Netdata gives more than you invest in it. A rare unicorn that obeys the Pareto rule.

— Eduard Porquet Mateu, TMB Barcelona

99% Downtime Reduction

Reduced website downtime by 99% and cloud bill by 30% using Netdata alerts.

— Falkland Islands Government

Real Savings
30% Cloud Cost Reduction

Optimized resource allocation based on Netdata alerts cut cloud spending by 30%.

— Falkland Islands Government

46% Cost Cut

Reduced monitoring staff by 67% while cutting operational costs by 46%.

— Codyas

Real Coverage
"Plugin for Everything"

Netdata has agent capacity or a plugin for everything, including Windows and Kubernetes.

— Eduard Porquet Mateu, TMB Barcelona

"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

Real Speed
Troubleshooting in 30 Seconds

From 2-3 minutes to 30 seconds—instant visibility into any node issue.

— Matthew Artist, Nodecraft

20% Downtime Reduction

20% less downtime and 40% budget optimization from out-of-the-box monitoring.

— Simon Beginn, LANCOM Systems

Pay per Node. Unlimited Everything Else.

One price per node. Unlimited metrics, logs, users, and retention. No per-GB surprises.

Free tier—forever
No metric limits or caps
Retention you control
Cancel anytime
> See pricing plans

What's Your Monitoring Really Costing You?

Most teams overpay by 40-60%. Let's find out why.

Expose hidden metric charges
Calculate tool consolidation
Customers report 30-67% savings
Results in under 60 seconds
> See what you're really paying

Your Infrastructure Is Unique. Let's Talk.

Because monitoring 10 nodes is different from monitoring 10,000.

On-prem & air-gapped deployment
Volume pricing & agreements
Architecture review for your scale
Compliance & security support
> Start a conversation

Monitoring That Sells Itself

Deploy in minutes. Impress clients in hours. Earn recurring revenue for years.

30-second live demos close deals
Zero config = zero support burden
Competitive margins & deal protection
Response in 48 hours
> Apply to partner

Per-Second Metrics at Homelab Prices

Same engine, same dashboards, same ML. Just priced for tinkerers.

Community: Free forever · 5 nodes · non-commercial
Homelab: $90/yr · unlimited nodes · fair usage
> Get the Homelab Plan

$1,000 Per Referral. Unlimited Referrals.

Your colleagues get 10% off. You get 10% commission. Everyone wins.

10% of subscriptions, up to $1,000 each
Track earnings inside Netdata Cloud
PayPal/Venmo payouts in 3-4 weeks
No caps, no complexity
> Get your referral link
Cost Proof
40% Budget Optimization

"Netdata's significant positive impact" — LANCOM Systems

Calculate Your Savings

Compare vs Datadog, Grafana, Dynatrace

Savings Proof
46% Cost Reduction

"Cut costs by 46%, staff by 67%" — Codyas

30% Cloud Bill Savings

"Reduced cloud bill by 30%" — Falkland Islands Gov

Enterprise Proof
"Better Than Combined Alternatives"

"Better observability with Netdata than combining other tools." — TMB Barcelona

Real Engineers, <24h Response

DPA, SLAs, on-prem, volume pricing

Why Partners Win
Demo Live Infrastructure

One command, 30 seconds, real data—no sandbox needed

Zero Tickets, High Margins

Auto-config + per-node pricing = predictable profit

Homelab Ready
Free Video Course

8-episode Netdata tutorial by LearnLinux.tv

76k+ GitHub Stars

3rd most starred monitoring project

Worth Recommending
Product That Delivers

Customers report 40-67% cost cuts, 99% downtime reduction

Zero Risk to Your Rep

Free tier lets them try before they buy

AI Support Assistant, Available 24/7

Nedi has access to all official documentation, source code, and resources. Ask any question about Netdata—responds in your language.

Deployment & configuration
Troubleshooting & sizing
Alerts & notifications
Evidence-based answers
> Ask Nedi now

Never Fight Fires Alone

Docs, community, and expert help—pick your path to resolution.

Learn.netdata.cloud docs
Discord, Forums, GitHub
Premium support available
> Get answers now

60 Seconds to First Dashboard

One command to install. Zero config. 850+ integrations documented.

Linux, Windows, K8s, Docker
Auto-discovers your stack
> Read our documentation

76,000+ Engineers Strong

615+ contributors. 1.5M daily downloads. One mission: simplify observability.

Per-Second. 90% Cheaper. Data Stays Home.

Side-by-side comparisons: costs, real-time granularity, and data sovereignty for every major tool.

See why teams switch from Datadog, Prometheus, Grafana, and more.

> Browse all comparisons
Edge-Native Observability, Born Open Source
Per-second visibility, ML on every metric, and data that never leaves your infrastructure.
Founded in 2016
615+ contributors worldwide
Remote-first, engineering-driven
Open source first
> Read our story
Promises We Publish—and Prove
12 principles backed by open code, independent validation, and measurable outcomes.
Open source, peer-reviewed
Zero config, instant value
Data sovereignty by design
Aligned pricing, no surprises
> See all 12 principles
Edge-Native, AI-Ready, 100% Open
76k+ stars. Full ML, AI, and automation—GPLv3+, not premium add-ons.
76,000+ GitHub stars
GPLv3+ licensed forever
ML on every metric, included
Zero vendor lock-in
> Explore our open source
Build Real-Time Observability for the World
Remote-first team shipping per-second monitoring with ML on every metric.
Remote-first, fully distributed
Open source (76k+ stars)
Challenging technical problems
Your code on millions of systems
> See open roles
Meet the Team Behind Netdata
Conferences, meetups, and tradeshows where you can see Netdata in action and talk to the engineers who build it.
Live demos and deep dives
Book 1-on-1 meetings
Talks and panel sessions
Event recaps and photos
> See all events
Talk to a Netdata Human in <24 Hours
Sales, partnerships, press, or professional services—real engineers, fast answers.
Discuss your observability needs
Pricing and volume discounts
Partnership opportunities
Media and press inquiries
> Book a conversation
Your Data. Your Rules.
On-prem data, cloud control plane, transparent terms.
Trust & Scale
76,000+ GitHub Stars

One of the most popular open-source monitoring projects

SOC 2 Type 2 Certified

Enterprise-grade security and compliance

Data Sovereignty

Your metrics stay on your infrastructure

Validated
University of Amsterdam

"Most energy-efficient monitoring solution" — ICSOC 2023, peer-reviewed

ADASTEC (Autonomous Driving)

"Doesn't miss alerts—mission-critical trust for safety software"

Community Stats
615+ Contributors

Global community improving monitoring for everyone

1.5M+ Downloads/Day

Trusted by teams worldwide

GPLv3+ Licensed

Free forever, fully open source agent

Why Join?
Remote-First

Work from anywhere, async-friendly culture

Impact at Scale

Your work helps millions of systems

$ guides / bind-dns / bind-dns-memory-growth-rss ▌

Operations Guides

BIND named RSS climbing: cache growth, allocator fragmentation, and real leaks

named RSS trending upward after warmup is the leading indicator before an OOM kill. The hard part is distinguishing three conditions that look identical on an RSS chart: cache-driven growth that will plateau under max-cache-size, allocator fragmentation that inflates RSS without a real leak, and genuine unbounded growth from a bug or misconfiguration.

RSS that climbs during a traffic burst or cache warmup and stays high is normal allocator behavior. BIND’s allocator (jemalloc on most builds) holds freed blocks for reuse rather than returning them to the OS via munmap. The operational question is not “is RSS high?” but “is RSS still growing, and how much runway remains?”

If RSS is growing linearly: time-to-OOM = (memory_limit - current_RSS) / growth_rate_per_day. A runway measured in weeks is a capacity planning problem. A runway measured in hours is an active incident.

What this means

Cache-driven growth (normal, bounded): The cache fills during warmup and after traffic bursts. RSS rises until the cache hits max-cache-size and LRU eviction begins (DeleteLRU increments). Growth should decelerate and plateau.

Allocator fragmentation (sticky but stable): The allocator holds freed blocks for reuse. RSS stays elevated even after cache entries expire or are evicted. Normal as long as RSS stabilizes. The gap between BIND’s internal accounting (InUse from the statistics channel) and OS-reported RSS (VmRSS from /proc) is the fragmentation overhead.

Unbounded growth (real problem): RSS grows monotonically without stabilizing. TreeMemInUse or HeapMemInUse trend upward without a corresponding increase in CacheNodes. This indicates a genuine leak, max-cache-size not effectively limiting cache memory, or non-cache consumers (RPZ datasets, zone data, ADB) growing outside the cache budget.

flowchart TD
    A["named RSS trending up"] --> B{"Still warming up?
uptime < 30 min"} B -->|Yes| C["Normal cache fill
Wait and re-check"] B -->|No| D{"RSS plateaued
or still growing?"} D -->|Plateaued| E["Normal: allocator
holds freed pages"] D -->|Still growing| F{"TreeMemInUse rising
without CacheNodes?"} F -->|Yes| G["Suspected leak
Check BIND version and CVEs"] F -->|No| H{"DeleteLRU
incrementing?"} H -->|Yes| I["Cache at capacity
Bound with max-cache-size"] H -->|No| J["Check non-cache memory:
RPZ, zone data, ADB"]

Common causes

CauseWhat it looks likeFirst thing to check
Cache warming (normal)RSS climbs for 30-60 min after restart, then plateaus. CacheHits rising, CacheMisses falling.Uptime. If under 1800 seconds, wait.
max-cache-size not configuredRSS grows toward 90% of physical RAM (the default for recursive views). No eviction pressure visible.named-checkconf -p | grep max-cache-size
max-cache-size ineffective on specific versionsRSS grows past the configured cap. DeleteLRU not incrementing despite cache exceeding limit.BIND version. See version-specific notes below.
Allocator fragmentationRSS high but stable. BIND internal InUse much lower than OS VmRSS. No monotonic growth trend.Compare internal vs OS memory over time.
CVE-2026-3104 memory leakUnbounded RSS growth on a recursive resolver. Assertion failure on shutdown or reload.BIND version: affects 9.20.0 through 9.20.20, 9.21.0 through 9.21.19. Fixed in 9.20.21 / 9.21.20.
jemalloc dirty-page glitchRSS grows slowly over days without stabilizing. No corresponding cache growth.MALLOC_CONF environment variable. BIND patch level.
RPZ datasetsRSS grows with RPZ zone load or after feed update. Memory is outside the cache budget.RPZ zone count and total entry count.
Zone dataRSS grows with authoritative zone count or large zone transfers. Not subject to max-cache-size.Zone count and total zone data size.

Quick checks

# Version and uptime
rndc status | head -5

# Current RSS in MB from /proc
awk '/VmRSS/{print $2/1024 " MB"}' /proc/$(pidof named)/status

# Peak RSS for context
awk '/VmPeak/{print $2/1024 " MB"}' /proc/$(pidof named)/status

# ISC-recommended RSS measurement: pmap Dirty column total (last line)
pmap -x $(pidof named) | tail -1

# Cache memory stats per view (TreeMemInUse, HeapMemInUse, CacheNodes, DeleteLRU)
# Adjust port to match your statistics-channels configuration (8053 and 8653 are common)
curl -s http://localhost:8653/json/v1/server | \
  python3 -c "import sys,json; d=json.load(sys.stdin); \
  [print(f'{v}: TreeMemInUse={cs.get(\"TreeMemInUse\",\"N/A\")} HeapMemInUse={cs.get(\"HeapMemInUse\",\"N/A\")} CacheNodes={cs.get(\"CacheNodes\",\"N/A\")} DeleteLRU={cs.get(\"DeleteLRU\",\"N/A\")}') \
  for v,vd in d.get('views',{}).items() if (cs:=vd.get('resolver',{}).get('cachestats',{}))]"

# Verify max-cache-size configuration
named-checkconf -p /etc/named.conf 2>/dev/null | grep -i "max-cache-size"

# Check for prior OOM kills of named
dmesg | grep -i "oom.*named\|named.*oom" | tail -10

# Check systemd restart count (OOM-kill loop indicator)
systemctl show named --property=NRestarts

# BIND internal memory context (total InUse vs Malloced)
curl -s http://localhost:8653/json/v1/mem | python3 -m json.tool | head -40

How to diagnose it

1. Confirm RSS is still growing. Sample RSS twice with a known interval and compute the daily rate. A single snapshot tells you nothing.

T1=$(date +%s); R1=$(awk '/VmRSS/{print $2}' /proc/$(pgrep -x named)/status); sleep 300
T2=$(date +%s); R2=$(awk '/VmRSS/{print $2}' /proc/$(pgrep -x named)/status)
# Daily growth rate in KB
echo "scale=0; ($R2 - $R1) * 86400 / ($T2 - $T1)" | bc

2. Rule out normal warmup. If uptime is under 1800 seconds, growth is cache warming. Wait and re-check. Cache hit rate should climb from near 0% toward baseline over 30-60 minutes.

3. Compare BIND internal memory with OS-reported RSS. If InUse from the statistics channel is stable or declining while VmRSS stays high, the gap is allocator fragmentation. If both climb in lockstep, the growth is real consumption.

4. Check whether the cache is at capacity and evicting. If DeleteLRU is incrementing, the cache has hit max-cache-size and is actively evicting. The cap is working. If TreeMemInUse exceeds the configured max-cache-size without DeleteLRU activity, the cap may not be effective for your BIND version.

5. Verify max-cache-size is configured. The default is 90% of physical memory for recursive views; this default dates to BIND 9.11.0. On multi-purpose servers, this default is dangerous. In chroot deployments where BIND cannot detect physical memory, the default is effectively unlimited. Set an explicit value.

6. Check BIND version against known memory bugs.

  • CVE-2026-3104: affects BIND 9.20.0 through 9.20.20 and 9.21.0 through 9.21.19. A crafted domain query causes unbounded RSS growth with no recovery. named also crashes with an assertion failure on shutdown or reload. Fixed in 9.20.21 and 9.21.20. If your resolver handles untrusted query traffic and is in an affected range, treat this as the primary suspect.
  • jemalloc dirty-page purging glitch: if a thread’s first allocator call is free() rather than malloc(), jemalloc’s dirty-page purging ticker is never initialized, causing RSS to creep upward over days. Fixed in BIND 9.16.29+, 9.18.3+, and all 9.20.x releases via a dummy allocation at thread start.

7. Identify non-cache memory consumers. Authoritative zone data, RPZ datasets, and the Address Database (ADB) all consume memory outside the max-cache-size budget. On a server with thousands of zones or large RPZ feeds, these can dwarf the cache.

Metrics and signals to monitor

SignalWhy it mattersWarning sign
VmRSS (/proc/$pid/status)Physical memory consumed. The gauge before OOM.Monotonic growth after warmup plateau.
TreeMemInUse / HeapMemInUseBIND’s internal cache memory accounting per view.Growing without CacheNodes increase indicates fragmentation or leak.
CacheNodesCurrent cache entry count. Correlates cache size with memory use.Flat or declining while TreeMemInUse grows means fragmentation.
DeleteLRULRU eviction count. Confirms cache is at capacity and cap is working.Zero while TreeMemInUse exceeds max-cache-size means cap is ineffective.
max-cache-sizeCaps cache memory. Default is 90% of RAM if unset.Not explicitly configured on shared hosts.
System available memoryHeadroom before OOM kill.Declining trend with no corresponding cache growth.
NRestarts (systemd)Restart count. OOM-kill loop indicator.Incrementing with OOM events in dmesg.

Fixes

Cache at capacity

If DeleteLRU is incrementing rapidly and cache hit ratio is declining, the cache is undersized for the working set. Options:

  • Raise max-cache-size. This directly increases the memory ceiling, so verify the host has headroom first. Apply with rndc reconfig after updating named.conf.
  • Add resolver capacity behind a load balancer to spread the query load and increase aggregate cache.
  • Accept the eviction rate if cache hit ratio is still above baseline.

See BIND cache eviction storms for the full pressure spiral pattern.

max-cache-size not configured or ineffective

Set an explicit value in named.conf:

options {
    max-cache-size 2g;
};

For recursive views, this caps the cache. For authoritative-only views where recursion no, BIND still applies a small 2 MB default cache limit even though that cache is normally unused. Apply changes with rndc reconfig.

Be aware that certain BIND 9.19 development snapshots turned max-cache-size into a no-op for overmem detection because the internal water-mark call was removed. The call was restored in 9.19.18, before the first 9.20 release. If you observe TreeMemInUse growing past your configured limit with no DeleteLRU activity, test on a patched version.

Allocator fragmentation (stable high RSS)

If RSS is high but stable, and BIND’s internal InUse is much lower than OS-reported VmRSS, this is fragmentation. Do not restart named to “fix” it. The restart gives you a temporary lower RSS but loses the warm cache, triggering a cache warming storm and elevated upstream load.

For the jemalloc dirty-page glitch on older BIND versions, setting MALLOC_CONF=dirty_decay_ms:0 forces immediate purging of dirty pages at a performance cost. This is a workaround. Upgrade to a version with the dummy-allocation fix instead.

Genuine memory leak

If RSS grows monotonically without stabilizing and you have ruled out cache growth, fragmentation, and non-cache consumers:

  1. Check your BIND version against CVE-2026-3104 (affects 9.20.0 through 9.20.20). Upgrade to 9.20.21 or later if in range.
  2. On BIND 9.20.6+, use rndc memprof on|off|dump to toggle or dump jemalloc profiling when the build and jemalloc configuration support it; otherwise it reports UNSUPPORTED.
  3. If running BIND 9.16.x with -M internal, upgrade if possible. BIND 9.18 reduced the internal allocator to a wrapper around the system allocator; BIND 9.20 removed the internal allocator and its separate counters.
  4. As a last resort, schedule rolling restarts during low-traffic windows to reset RSS. Track the growth rate to estimate how frequently restarts are needed.

Non-cache memory (RPZ, zones, ADB)

Zone data and RPZ datasets are not bounded by max-cache-size. If these are the source of growth:

  • Monitor RPZ feed sizes independently. A threat intelligence feed that doubles in size doubles its memory footprint.
  • For authoritative servers with many zones, account for zone data separately in memory planning.
  • The ADB tracks reachability and RTT to upstream authoritative servers. On a resolver resolving many unique upstream targets, ADB memory can be significant.

Prevention

Set max-cache-size explicitly on every recursive resolver. The default of 90% of physical RAM is appropriate only on dedicated DNS hosts. On shared infrastructure, use a fixed value that leaves room for the OS and other processes. Target 70% of available memory as the ceiling for total named RSS.

Track RSS as a trend, not a threshold. A single RSS value is meaningless. Track the daily growth rate after warmup. If it is nonzero and sustained, compute runway: (memory_limit - current_RSS) / growth_rate_per_day. Alert when runway drops below 14 days.

Keep BIND patched. Memory leaks in BIND are typically version-specific bugs. CVE-2026-3104 and the jemalloc glitch are both fixed in recent releases. Running a supported Extended Support Version with current patches is the simplest defense.

Separate recursive and authoritative roles. Mixed-role servers make memory diagnosis harder because cache growth and zone data growth are indistinguishable in aggregate RSS. Cache memory is bounded by max-cache-size; zone data is not.

Avoid polling the catch-all /json statistics endpoint too frequently. On hosts with many CPUs, a single fetch of the full JSON statistics can serialize tens of thousands of task objects and trigger a transient memory spike. Use granular endpoints (/json/v1/server, /json/v1/mem) instead.

How Netdata helps

Netdata’s BIND collector gathers per-view TreeMemInUse, HeapMemInUse, CacheNodes, DeleteLRU, and DeleteTTL at per-second resolution, correlated on the same timeline as process RSS (VmRSS), available system memory, and cgroup limits. This makes the fragmentation gap visible as the spread between internal memory accounting and OS-reported RSS.

The OOM kill detector surfaces kernel log events alongside the RSS trend that preceded them. After a restart, the cold-cache signature (zero CacheHits, elevated outbound query rate) appears on the same chart as the RSS reset.

ML-based anomaly detection on the RSS growth rate can flag a changing growth pattern before it crosses a fixed threshold, providing earlier runway warning than a static alert.