The only agent that thinks for itself

Autonomous Monitoring with self-learning AI built-in, operating independently across your entire stack.

Unlimited Metrics & Logs
Machine learning & MCP
5% CPU, 150MB RAM
3GB disk, >1 year retention
800+ integrations, zero config
Dashboards, alerts out of the box
> Discover Netdata Agents

Centralized metrics streaming and storage

Aggregate metrics from multiple agents into centralized Parent nodes for unified monitoring across your infrastructure.

Stream from unlimited agents
Long-term data retention
High availability clustering
Data replication & backup
Scalable architecture
Enterprise-grade security
> Learn about Parents

Fully managed cloud platform

Access your monitoring data from anywhere with our SaaS platform. No infrastructure to manage, automatic updates, and global availability.

Zero infrastructure management
99.9% uptime SLA
Global data centers
Automatic updates & patches
Enterprise SSO & RBAC
SOC2 & ISO certified
> Explore Netdata Cloud

Deploy Netdata Cloud in your infrastructure

Run the full Netdata Cloud platform on-premises for complete data sovereignty and compliance with your security policies.

Complete data sovereignty
Air-gapped deployment
Custom compliance controls
Private network integration
Dedicated support team
Kubernetes & Docker support
> Learn about Cloud On-Premises

Powerful, intuitive monitoring interface

Modern, responsive UI built for real-time troubleshooting with customizable dashboards and advanced visualization capabilities.

Real-time chart updates
Customizable dashboards
Dark & light themes
Advanced filtering & search
Responsive on all devices
Collaboration features
> Explore Netdata UI

Monitor on the go

Native iOS and Android apps bring full monitoring capabilities to your mobile device with real-time alerts and notifications.

iOS & Android apps
Push notifications
Touch-optimized interface
Offline data access
Biometric authentication
Widget support
> Download apps

The future of infrastructure observability

See our strategic direction across AI-native observability, full-stack signals, operational intelligence, and enterprise platform maturity.

AI-native observability
Full-stack signal coverage
Operational intelligence
Enterprise platform maturity
Agent releases every 6 weeks
Cloud continuous delivery
> Explore Product Roadmap

Best energy efficiency

True real-time per-second

100% automated zero config

Centralized observability

Multi-year retention

High availability built-in

Zero maintenance

Always up-to-date

Enterprise security

Complete data control

Air-gap ready

Compliance certified

Millisecond responsiveness

Infinite zoom & pan

Works on any device

Native performance

Instant alerts

Monitor anywhere

AI-native observability

Continuous delivery

Open source foundation

80% Faster Incident Resolution

AI-powered troubleshooting from detection, to root cause and blast radius identification, to reporting.

True Real-Time and Simple, even at Scale

Linearly and infinitely scalable full-stack observability, that can be deployed even mid-crisis.

90% Cost Reduction, Full Fidelity

Instead of centralizing the data, Netdata distributes the code, eliminating pipelines and complexity.

See and Map Your Entire Network

Live topology, flow analytics, and SNMP device and trap monitoring — unified with your full-stack observability.

Control Without Surrender

SOC 2 Type 2 certified with every metric kept on your infrastructure.

Integrations

800+ collectors and notification channels, auto-discovered and ready out of the box.

800+ data collectors
Auto-discovery & zero config
Cloud, infra, app protocols
Notifications out of the box
> Explore integrations
Real Results
46% Cost Reduction

Reduced monitoring costs by 46% while cutting staff overhead by 67%.

— Leonardo Antunez, Codyas

Zero Pipeline

No data shipping. No central storage costs. Query at the edge.

From Our Users
"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

No Query Language

Point-and-click troubleshooting. No PromQL, no LogQL, no learning curve.

Enterprise Ready
67% Less Staff, 46% Cost Cut

Enterprise efficiency without enterprise complexity—real ROI from day one.

— Leonardo Antunez, Codyas

SOC 2 Type 2 Certified

Zero data egress. Only metadata reaches the cloud. Your metrics stay on your infrastructure.

Full Coverage
800+ Collectors

Auto-discovered and configured. No manual setup required.

Any Notification Channel

Slack, PagerDuty, Teams, email, webhooks—all built-in.

Built for the People Who Get Paged

Because 3am alerts deserve instant answers, not hour-long hunts.

Every Industry Has Rules. We Master Them.

See how healthcare, finance, and government teams cut monitoring costs 90% while staying audit-ready.

Monitor Any Technology. Configure Nothing.

Install the agent. It already knows your stack.
From Our Users
"A Rare Unicorn"

Netdata gives more than you invest in it. A rare unicorn that obeys the Pareto rule.

— Eduard Porquet Mateu, TMB Barcelona

99% Downtime Reduction

Reduced website downtime by 99% and cloud bill by 30% using Netdata alerts.

— Falkland Islands Government

Real Savings
30% Cloud Cost Reduction

Optimized resource allocation based on Netdata alerts cut cloud spending by 30%.

— Falkland Islands Government

46% Cost Cut

Reduced monitoring staff by 67% while cutting operational costs by 46%.

— Codyas

Real Coverage
"Plugin for Everything"

Netdata has agent capacity or a plugin for everything, including Windows and Kubernetes.

— Eduard Porquet Mateu, TMB Barcelona

"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

Real Speed
Troubleshooting in 30 Seconds

From 2-3 minutes to 30 seconds—instant visibility into any node issue.

— Matthew Artist, Nodecraft

20% Downtime Reduction

20% less downtime and 40% budget optimization from out-of-the-box monitoring.

— Simon Beginn, LANCOM Systems

Pay per Node. Unlimited Everything Else.

One price per node. Unlimited metrics, logs, users, and retention. No per-GB surprises.

Free tier—forever
No metric limits or caps
Retention you control
Cancel anytime
> See pricing plans

What's Your Monitoring Really Costing You?

Most teams overpay by 40-60%. Let's find out why.

Expose hidden metric charges
Calculate tool consolidation
Customers report 30-67% savings
Results in under 60 seconds
> See what you're really paying

Your Infrastructure Is Unique. Let's Talk.

Because monitoring 10 nodes is different from monitoring 10,000.

On-prem & air-gapped deployment
Volume pricing & agreements
Architecture review for your scale
Compliance & security support
> Start a conversation

Monitoring That Sells Itself

Deploy in minutes. Impress clients in hours. Earn recurring revenue for years.

30-second live demos close deals
Zero config = zero support burden
Competitive margins & deal protection
Response in 48 hours
> Apply to partner

Per-Second Metrics at Homelab Prices

Same engine, same dashboards, same ML. Just priced for tinkerers.

Community: Free forever · 5 nodes · non-commercial
Homelab: $90/yr · unlimited nodes · fair usage
> Get the Homelab Plan

$1,000 Per Referral. Unlimited Referrals.

Your colleagues get 10% off. You get 10% commission. Everyone wins.

10% of subscriptions, up to $1,000 each
Track earnings inside Netdata Cloud
PayPal/Venmo payouts in 3-4 weeks
No caps, no complexity
> Get your referral link
Cost Proof
40% Budget Optimization

"Netdata's significant positive impact" — LANCOM Systems

Calculate Your Savings

Compare vs Datadog, Grafana, Dynatrace

Savings Proof
46% Cost Reduction

"Cut costs by 46%, staff by 67%" — Codyas

30% Cloud Bill Savings

"Reduced cloud bill by 30%" — Falkland Islands Gov

Enterprise Proof
"Better Than Combined Alternatives"

"Better observability with Netdata than combining other tools." — TMB Barcelona

Real Engineers, <24h Response

DPA, SLAs, on-prem, volume pricing

Why Partners Win
Demo Live Infrastructure

One command, 30 seconds, real data—no sandbox needed

Zero Tickets, High Margins

Auto-config + per-node pricing = predictable profit

Homelab Ready
Free Video Course

8-episode Netdata tutorial by LearnLinux.tv

76k+ GitHub Stars

3rd most starred monitoring project

Worth Recommending
Product That Delivers

Customers report 40-67% cost cuts, 99% downtime reduction

Zero Risk to Your Rep

Free tier lets them try before they buy

AI Support Assistant, Available 24/7

Nedi has access to all official documentation, source code, and resources. Ask any question about Netdata—responds in your language.

Deployment & configuration
Troubleshooting & sizing
Alerts & notifications
Evidence-based answers
> Ask Nedi now

Never Fight Fires Alone

Docs, community, and expert help—pick your path to resolution.

Learn.netdata.cloud docs
Discord, Forums, GitHub
Premium support available
> Get answers now

60 Seconds to First Dashboard

One command to install. Zero config. 850+ integrations documented.

Linux, Windows, K8s, Docker
Auto-discovers your stack
> Read our documentation

76,000+ Engineers Strong

615+ contributors. 1.5M daily downloads. One mission: simplify observability.

Per-Second. 90% Cheaper. Data Stays Home.

Side-by-side comparisons: costs, real-time granularity, and data sovereignty for every major tool.

See why teams switch from Datadog, Prometheus, Grafana, and more.

> Browse all comparisons
Edge-Native Observability, Born Open Source
Per-second visibility, ML on every metric, and data that never leaves your infrastructure.
Founded in 2016
615+ contributors worldwide
Remote-first, engineering-driven
Open source first
> Read our story
Promises We Publish—and Prove
12 principles backed by open code, independent validation, and measurable outcomes.
Open source, peer-reviewed
Zero config, instant value
Data sovereignty by design
Aligned pricing, no surprises
> See all 12 principles
Edge-Native, AI-Ready, 100% Open
76k+ stars. Full ML, AI, and automation—GPLv3+, not premium add-ons.
76,000+ GitHub stars
GPLv3+ licensed forever
ML on every metric, included
Zero vendor lock-in
> Explore our open source
Build Real-Time Observability for the World
Remote-first team shipping per-second monitoring with ML on every metric.
Remote-first, fully distributed
Open source (76k+ stars)
Challenging technical problems
Your code on millions of systems
> See open roles
Meet the Team Behind Netdata
Conferences, meetups, and tradeshows where you can see Netdata in action and talk to the engineers who build it.
Live demos and deep dives
Book 1-on-1 meetings
Talks and panel sessions
Event recaps and photos
> See all events
Talk to a Netdata Human in <24 Hours
Sales, partnerships, press, or professional services—real engineers, fast answers.
Discuss your observability needs
Pricing and volume discounts
Partnership opportunities
Media and press inquiries
> Book a conversation
Your Data. Your Rules.
On-prem data, cloud control plane, transparent terms.
Trust & Scale
76,000+ GitHub Stars

One of the most popular open-source monitoring projects

SOC 2 Type 2 Certified

Enterprise-grade security and compliance

Data Sovereignty

Your metrics stay on your infrastructure

Validated
University of Amsterdam

"Most energy-efficient monitoring solution" — ICSOC 2023, peer-reviewed

ADASTEC (Autonomous Driving)

"Doesn't miss alerts—mission-critical trust for safety software"

Community Stats
615+ Contributors

Global community improving monitoring for everyone

1.5M+ Downloads/Day

Trusted by teams worldwide

GPLv3+ Licensed

Free forever, fully open source agent

Why Join?
Remote-First

Work from anywhere, async-friendly culture

Impact at Scale

Your work helps millions of systems

$ guides / bind-dns / bind-dns-random-subdomain-attack ▌

Operations Guides

BIND random subdomain (water torture) attack: an NXDOMAIN flood that bypasses the cache

Your BIND resolver is up, the process is running, UDP and TCP both answer on port 53. But clients across the network are experiencing slow DNS or outright SERVFAIL. NXDOMAIN rate has spiked to several times baseline. Cache hit ratio is collapsing. Recursive clients are climbing toward the hard limit. Upstream query rate has ballooned to approach or exceed the inbound rate.

This is a random subdomain attack, also called water torture or PRSD (Persistent Random Subdomain). Attackers flood your resolver with queries for randomized subdomains of a real domain: a1b2c3.victim.com, x7y8z9.victim.com, q2w3e4.victim.com. Every name is unique, so the cache cannot help. Each query is a miss, each miss forces recursion, and each recursion consumes a recursive-client slot and CPU. The victim’s authoritative server is hammered, but the real damage is collateral: your resolver becomes so saturated that resolution for every client, for every domain, degrades.

The distinguishing feature is high query-name cardinality concentrated on one parent domain with near-zero repetition. Legitimate traffic repeats names (clients ask for google.com thousands of times). Attack traffic does not.

What this means

A warm BIND cache answers 80-95% of queries from memory in under a millisecond. The water torture attack defeats the cache entirely by ensuring no two queries ever match.

Each unique random subdomain triggers the full recursive pipeline: cache lookup (miss), iterative query to the victim’s authoritative nameservers, wait for NXDOMAIN response. Each in-flight query holds a slot in the recursive-clients table (default limit 1000, soft quota at 900). The victim’s authoritative server is also under attack, so responses may be slow or time out. Default resolver-query-timeout is 10 seconds. If the victim’s authoritative servers are unresponsive, a single failed query holds a recursive-client slot for the full window.

As slots fill, fewer are available for legitimate queries. At the soft quota (90% of a 1000 limit), BIND logs a rate-limited warning and reclaims its oldest slot; the incoming query is still served. At the hard limit, BIND also reclaims its oldest slot and the new recursive query receives SERVFAIL. The attack on one domain has now become a service-wide outage.

flowchart TD
    A[Random unique queries for victim.com subdomains] --> B[Every name is a cache miss]
    B --> C[Forced recursion to upstream NS]
    C --> D[RecursClients slots consumed]
    D --> E[NXDOMAIN rate and CPU spike]
    D --> F{recursive-clients near limit?}
    F -->|No| G[Degraded resolution for all clients]
    F -->|Yes| H[SERVFAIL for all recursive queries]

One detail that makes this attack particularly effective against BIND: clients-per-query and max-clients-per-query (BIND’s deduplication for concurrent identical recursive queries) provide no protection. These controls collapse multiple clients asking for the same name into a single upstream fetch. When every query name is unique, there is nothing to deduplicate.

Common causes

CauseWhat it looks likeFirst thing to check
Botnet-driven water tortureHigh NXDOMAIN rate concentrated on one parent domain, near-zero query name repetitionQuery log cardinality analysis or rndc dumpdb -cache
Compromised IoT or internal devicesAttack traffic originates from your own subnetsSource IP distribution in query logs
Misconfigured application retry loopHigh query rate from one or few sources, names may be random or semi-randomSource IP concentration and query pattern

The first two are the classic pattern. The third is less common but mimics the same signal profile. In all cases, the damage path is identical: cache misses cascade into recursive-client exhaustion.

Quick checks

These commands assume BIND’s statistics channel is enabled on port 8653. Adjust the port to match your statistics-channels configuration. All commands are read-only.

# Check NXDOMAIN rate relative to other response codes
curl -s http://localhost:8653/json/v1/server | \
  python3 -c "import sys,json; d=json.load(sys.stdin); \
  [print(f'{k}: {v}') for k,v in sorted(d.get('nsstats',{}).items()) if k.startswith('Qry')]"

# Check recursive clients as a gauge (absolute value, not a counter)
curl -s http://localhost:8653/json/v1/server | \
  python3 -c "import sys,json; d=json.load(sys.stdin); \
  print('RecursClients:', d.get('nsstats',{}).get('RecursClients', 'N/A'))"

# Check cache hit ratio per view
curl -s http://localhost:8653/json/v1/server | \
  python3 -c "import sys,json; d=json.load(sys.stdin); \
  [print(f'{v}: Hits={cs.get(\"CacheHits\",0)} Misses={cs.get(\"CacheMisses\",0)} Ratio={cs.get(\"CacheHits\",0)/(cs.get(\"CacheHits\",0)+cs.get(\"CacheMisses\",1))*100:.1f}%') \
  for v,vd in d.get('views',{}).items() if (cs:=vd.get('resolver',{}).get('cachestats',{}))]"

# See current recursive queries and active iterative-fetch domains
rndc recursing

# Check RRL activity (only present if rate-limit is configured)
curl -s http://localhost:8653/json/v1/server | \
  python3 -c "import sys,json; d=json.load(sys.stdin); ns=d.get('nsstats',{}); \
  print('RateDropped:', ns.get('RateDropped',0), 'RateSlipped:', ns.get('RateSlipped',0))"

# Check CPU utilization for named process
pidstat -p $(pidof named) 1 5

If RecursClients is above 900 (90% of the default 1000 limit), NXDOMAIN is dominating the response code distribution, and cache hit ratio has dropped well below baseline, you are under attack or experiencing a similar high-cardinality query pattern.

How to diagnose it

  1. Identify the targeted domain. Dump the cache and analyze query name concentration:

    rndc dumpdb -cache
    # Dumps to the working directory (default: /var/named or as configured).
    # Can be memory and I/O intensive on large caches. Analyze the dump
    # for the most common parent domain.
    

    Alternatively, if query logging is available (enable it briefly; it degrades performance at high QPS), aggregate by parent domain:

    awk '{print $NF}' /var/log/named/queries.log | rev | cut -d. -f1-2 | rev | sort | uniq -c | sort -rn | head -10
    
  2. Confirm the attack pattern. The signature is near-zero repetition among query names. Legitimate high-NXDOMAIN sources produce different distributions: Windows DNS suffix search lists generate predictable suffix-appended names, and DGA malware produces many different parent domains. Water torture concentrates randomness under one parent.

  3. Assess collateral damage. Check whether RecursClients is approaching the limit and whether SERVFAIL is rising for unrelated domains. If both are true, the attack is already degrading resolution for all clients.

  4. Check upstream victim status. Use rndc recursing to identify current recursive queries and active iterative-fetch domains. If the victim is also slow (because they are under the same attack), timeout duration amplifies slot consumption.

  5. Check kernel UDP drops. The attack volume may exceed the kernel’s ability to buffer packets:

    cat /proc/net/snmp | grep Udp
    # Watch UdpRcvbufErrors for non-zero increments
    

    These drops are invisible to BIND. If present, your resolver is losing queries before it processes them.

Metrics and signals to monitor

SignalWhy it mattersWarning sign
QryNXDOMAIN rateNXDOMAIN is the dominant response for random subdomain queriesSustained NXDOMAIN at 3x or more above baseline
Cache hit ratio (CacheHits / CacheMisses)The cache is useless when every query is uniqueDrop of 15+ percentage points from rolling baseline
RecursClients (gauge)Each forced recursion holds a slot; exhaustion means SERVFAIL for all queriesAbove 50% of recursive-clients limit sustained
NumFetch (per view)Per-view recursive pressure in split-horizon setupsSustained upward trend
Outbound query rateApproaching inbound rate means cache provides no valueOutbound/inbound ratio above 50% on a warm resolver
RateDropped / RateSlippedRRL is actively limiting (only if rate-limit configured)Non-zero sustained values
RPZRewritesRPZ policy is intercepting queries (only if RPZ configured)Spike correlating with NXDOMAIN drop after RPZ applied
UdpRcvbufErrorsKernel is dropping packets before BIND sees themAny sustained non-zero rate
CPU for namedQuery processing and recursion consume CPU per unique querySustained above 70% of available cores

Fixes

Apply RPZ to refuse the targeted domain

The fastest immediate mitigation. Configure a Response Policy Zone that returns NXDOMAIN for the victim domain before recursion is triggered. RPZ policy is evaluated early in the query pipeline, so the query never reaches the recursive fetch stage.

In the RPZ zone file, encode an NXDOMAIN policy with a CNAME . record for the target domain:

victim.com     CNAME .

Apply it with rndc reload of the RPZ zone. Confirm it is working by checking that RPZRewrites increases and NXDOMAIN responses for the targeted domain now come from policy rather than from upstream recursion.

Tradeoff: RPZ requires knowing the target domain. If the attacker shifts to a new domain, you must update the RPZ rule. RPZ datasets also consume memory, so factor that into capacity planning.

Enable rate-limit (RRL)

Response Rate Limiting throttles outgoing responses. RRL is built into BIND since version 9.9 and requires no compile flag. Configure it in named.conf:

rate-limit {
    responses-per-second 100;
};

The default responses-per-second is 0, meaning no limit. You must set it explicitly. The nxdomains-per-second option defaults to responses-per-second, so setting the parent value also limits NXDOMAIN responses.

RRL distinguishes between RateDropped (responses silently dropped) and RateSlipped (responses sent truncated, forcing TCP retry). Monitor both after enabling. Slipped responses cause TCP retry, which can create FD pressure if TCP capacity is also constrained.

Tradeoff: Overly aggressive RRL settings can throttle legitimate clients. Tune based on your normal traffic baseline.

Configure fetches-per-zone

fetches-per-zone caps the number of simultaneous iterative queries BIND will send for any single domain. This directly limits how many recursive slots a water torture attack can consume against the targeted domain.

options {
    fetches-per-zone 100;
};

The default is 0 (disabled). When the limit is exceeded, the default action is drop (the query is silently dropped). You can also specify fail to return SERVFAIL instead. This option was introduced in BIND 9.10.3.

Tradeoff: Setting this too low can affect legitimate high-traffic domains that genuinely generate many concurrent recursive queries. Test with your normal traffic patterns before applying in production.

Configure fetches-per-server

This limits simultaneous queries to any single upstream nameserver. The default is 0 (disabled). BIND can also adaptively adjust this quota downward based on upstream responsiveness via fetch-quota-params.

options {
    fetches-per-server 100;
};

The default action when the limit is exceeded is fail (returns SERVFAIL). If the victim’s authoritative servers are slow because they are also under attack, this limit prevents your resolver from piling up queries against them. But legitimate queries to those servers are also capped.

Flush cache for the targeted domain

If stuck cache entries are compounding the issue:

rndc flushtree victim.com

rndc flushtree clears the specified domain and its subdomains from the view DNS cache, nameserver address database, bad-server cache, and SERVFAIL cache; data outside that name tree is unaffected.

Do not rely on stale NXDOMAIN caching

Since BIND 9.18.5 and 9.19.3 (change GL #3386), NXDOMAIN records are no longer retained past their normal negative cache TTL, even when stale-cache-enable yes is configured.

This is a deliberate defense against memory exhaustion during water torture attacks. It also means you cannot count on stale NXDOMAIN entries to absorb repeat queries. The attack generates unique names regardless, so stale caching would not help even if it were available.

Prevention

  • Pre-deploy RPZ infrastructure. Having RPZ configured and ready means you can add a refuse rule in seconds during an attack, without building the infrastructure under pressure.
  • Pre-configure fetches-per-zone. Set a reasonable limit (for example, 100) before an attack occurs. The default of 0 means no protection.
  • Monitor query name cardinality. This is a Level 4 maturity signal. Low entropy (repeated names) is normal. High entropy (random strings concentrated under one parent) indicates water torture. Sampling query logs periodically can detect this without the I/O cost of continuous query logging.
  • Monitor NXDOMAIN as a ratio of total responses. A sustained spike above 3x baseline warrants investigation. The NXDOMAIN spike guide covers broader context on NXDOMAIN pattern analysis including DGA malware and Windows suffix search lists.
  • Size recursive-clients appropriately. The default of 1000 is too low for busy resolvers. Increasing it without corresponding FD and memory capacity just delays the cliff. Scale it with available resources.
  • Keep BIND patched. If using nxdomain-redirect as part of your mitigation strategy, ensure you are on BIND 9.18.24 or later: CVE-2023-5517 can make named exit with an assertion failure when nxdomain-redirect is enabled and a matching PTR query is received. CVE-2023-4408 is a separate excessive-CPU flaw in message parsing, also fixed in 9.18.24.

How Netdata helps

The attack produces a multi-signal signature that unfolds over seconds, not minutes. Per-second metric collection catches this before thresholds trip:

  • NXDOMAIN rate correlation. Netdata surfaces QryNXDOMAIN alongside other response codes in real time. A sudden spike in NXDOMAIN share while NOERROR stays flat is the leading edge of the attack.
  • Cache hit ratio trend. Per-second cache hit and miss collection shows the ratio collapse within seconds, not after a 5-minute polling interval.
  • RecursClients as a gauge. Netdata collects RecursClients as an absolute value (not a counter), so you see the climb toward the limit in real time. Combined with the recursive-clients config value, the percentage-of-limit view is immediately available.
  • Outbound query rate. The outbound-to-inbound ratio is the clearest signal that the cache has stopped providing value. Per-second resolution catches this before the recursive-client limit is reached.
  • Anomaly detection. Netdata’s ML flags the sudden shift in response code distribution and cache behavior before explicit thresholds are crossed.