The only agent that thinks for itself

Autonomous Monitoring with self-learning AI built-in, operating independently across your entire stack.

Unlimited Metrics & Logs
Machine learning & MCP
5% CPU, 150MB RAM
3GB disk, >1 year retention
800+ integrations, zero config
Dashboards, alerts out of the box
> Discover Netdata Agents

Centralized metrics streaming and storage

Aggregate metrics from multiple agents into centralized Parent nodes for unified monitoring across your infrastructure.

Stream from unlimited agents
Long-term data retention
High availability clustering
Data replication & backup
Scalable architecture
Enterprise-grade security
> Learn about Parents

Fully managed cloud platform

Access your monitoring data from anywhere with our SaaS platform. No infrastructure to manage, automatic updates, and global availability.

Zero infrastructure management
99.9% uptime SLA
Global data centers
Automatic updates & patches
Enterprise SSO & RBAC
SOC2 & ISO certified
> Explore Netdata Cloud

Deploy Netdata Cloud in your infrastructure

Run the full Netdata Cloud platform on-premises for complete data sovereignty and compliance with your security policies.

Complete data sovereignty
Air-gapped deployment
Custom compliance controls
Private network integration
Dedicated support team
Kubernetes & Docker support
> Learn about Cloud On-Premises

Powerful, intuitive monitoring interface

Modern, responsive UI built for real-time troubleshooting with customizable dashboards and advanced visualization capabilities.

Real-time chart updates
Customizable dashboards
Dark & light themes
Advanced filtering & search
Responsive on all devices
Collaboration features
> Explore Netdata UI

Monitor on the go

Native iOS and Android apps bring full monitoring capabilities to your mobile device with real-time alerts and notifications.

iOS & Android apps
Push notifications
Touch-optimized interface
Offline data access
Biometric authentication
Widget support
> Download apps

The future of infrastructure observability

See our strategic direction across AI-native observability, full-stack signals, operational intelligence, and enterprise platform maturity.

AI-native observability
Full-stack signal coverage
Operational intelligence
Enterprise platform maturity
Agent releases every 6 weeks
Cloud continuous delivery
> Explore Product Roadmap

Best energy efficiency

True real-time per-second

100% automated zero config

Centralized observability

Multi-year retention

High availability built-in

Zero maintenance

Always up-to-date

Enterprise security

Complete data control

Air-gap ready

Compliance certified

Millisecond responsiveness

Infinite zoom & pan

Works on any device

Native performance

Instant alerts

Monitor anywhere

AI-native observability

Continuous delivery

Open source foundation

80% Faster Incident Resolution

AI-powered troubleshooting from detection, to root cause and blast radius identification, to reporting.

True Real-Time and Simple, even at Scale

Linearly and infinitely scalable full-stack observability, that can be deployed even mid-crisis.

90% Cost Reduction, Full Fidelity

Instead of centralizing the data, Netdata distributes the code, eliminating pipelines and complexity.

See and Map Your Entire Network

Live topology, flow analytics, and SNMP device and trap monitoring — unified with your full-stack observability.

Control Without Surrender

SOC 2 Type 2 certified with every metric kept on your infrastructure.

Integrations

800+ collectors and notification channels, auto-discovered and ready out of the box.

800+ data collectors
Auto-discovery & zero config
Cloud, infra, app protocols
Notifications out of the box
> Explore integrations
Real Results
46% Cost Reduction

Reduced monitoring costs by 46% while cutting staff overhead by 67%.

— Leonardo Antunez, Codyas

Zero Pipeline

No data shipping. No central storage costs. Query at the edge.

From Our Users
"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

No Query Language

Point-and-click troubleshooting. No PromQL, no LogQL, no learning curve.

Enterprise Ready
67% Less Staff, 46% Cost Cut

Enterprise efficiency without enterprise complexity—real ROI from day one.

— Leonardo Antunez, Codyas

SOC 2 Type 2 Certified

Zero data egress. Only metadata reaches the cloud. Your metrics stay on your infrastructure.

Full Coverage
800+ Collectors

Auto-discovered and configured. No manual setup required.

Any Notification Channel

Slack, PagerDuty, Teams, email, webhooks—all built-in.

Built for the People Who Get Paged

Because 3am alerts deserve instant answers, not hour-long hunts.

Every Industry Has Rules. We Master Them.

See how healthcare, finance, and government teams cut monitoring costs 90% while staying audit-ready.

Monitor Any Technology. Configure Nothing.

Install the agent. It already knows your stack.
From Our Users
"A Rare Unicorn"

Netdata gives more than you invest in it. A rare unicorn that obeys the Pareto rule.

— Eduard Porquet Mateu, TMB Barcelona

99% Downtime Reduction

Reduced website downtime by 99% and cloud bill by 30% using Netdata alerts.

— Falkland Islands Government

Real Savings
30% Cloud Cost Reduction

Optimized resource allocation based on Netdata alerts cut cloud spending by 30%.

— Falkland Islands Government

46% Cost Cut

Reduced monitoring staff by 67% while cutting operational costs by 46%.

— Codyas

Real Coverage
"Plugin for Everything"

Netdata has agent capacity or a plugin for everything, including Windows and Kubernetes.

— Eduard Porquet Mateu, TMB Barcelona

"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

Real Speed
Troubleshooting in 30 Seconds

From 2-3 minutes to 30 seconds—instant visibility into any node issue.

— Matthew Artist, Nodecraft

20% Downtime Reduction

20% less downtime and 40% budget optimization from out-of-the-box monitoring.

— Simon Beginn, LANCOM Systems

Pay per Node. Unlimited Everything Else.

One price per node. Unlimited metrics, logs, users, and retention. No per-GB surprises.

Free tier—forever
No metric limits or caps
Retention you control
Cancel anytime
> See pricing plans

What's Your Monitoring Really Costing You?

Most teams overpay by 40-60%. Let's find out why.

Expose hidden metric charges
Calculate tool consolidation
Customers report 30-67% savings
Results in under 60 seconds
> See what you're really paying

Your Infrastructure Is Unique. Let's Talk.

Because monitoring 10 nodes is different from monitoring 10,000.

On-prem & air-gapped deployment
Volume pricing & agreements
Architecture review for your scale
Compliance & security support
> Start a conversation

Monitoring That Sells Itself

Deploy in minutes. Impress clients in hours. Earn recurring revenue for years.

30-second live demos close deals
Zero config = zero support burden
Competitive margins & deal protection
Response in 48 hours
> Apply to partner

Per-Second Metrics at Homelab Prices

Same engine, same dashboards, same ML. Just priced for tinkerers.

Community: Free forever · 5 nodes · non-commercial
Homelab: $90/yr · unlimited nodes · fair usage
> Get the Homelab Plan

$1,000 Per Referral. Unlimited Referrals.

Your colleagues get 10% off. You get 10% commission. Everyone wins.

10% of subscriptions, up to $1,000 each
Track earnings inside Netdata Cloud
PayPal/Venmo payouts in 3-4 weeks
No caps, no complexity
> Get your referral link
Cost Proof
40% Budget Optimization

"Netdata's significant positive impact" — LANCOM Systems

Calculate Your Savings

Compare vs Datadog, Grafana, Dynatrace

Savings Proof
46% Cost Reduction

"Cut costs by 46%, staff by 67%" — Codyas

30% Cloud Bill Savings

"Reduced cloud bill by 30%" — Falkland Islands Gov

Enterprise Proof
"Better Than Combined Alternatives"

"Better observability with Netdata than combining other tools." — TMB Barcelona

Real Engineers, <24h Response

DPA, SLAs, on-prem, volume pricing

Why Partners Win
Demo Live Infrastructure

One command, 30 seconds, real data—no sandbox needed

Zero Tickets, High Margins

Auto-config + per-node pricing = predictable profit

Homelab Ready
Free Video Course

8-episode Netdata tutorial by LearnLinux.tv

76k+ GitHub Stars

3rd most starred monitoring project

Worth Recommending
Product That Delivers

Customers report 40-67% cost cuts, 99% downtime reduction

Zero Risk to Your Rep

Free tier lets them try before they buy

AI Support Assistant, Available 24/7

Nedi has access to all official documentation, source code, and resources. Ask any question about Netdata—responds in your language.

Deployment & configuration
Troubleshooting & sizing
Alerts & notifications
Evidence-based answers
> Ask Nedi now

Never Fight Fires Alone

Docs, community, and expert help—pick your path to resolution.

Learn.netdata.cloud docs
Discord, Forums, GitHub
Premium support available
> Get answers now

60 Seconds to First Dashboard

One command to install. Zero config. 850+ integrations documented.

Linux, Windows, K8s, Docker
Auto-discovers your stack
> Read our documentation

76,000+ Engineers Strong

615+ contributors. 1.5M daily downloads. One mission: simplify observability.

Per-Second. 90% Cheaper. Data Stays Home.

Side-by-side comparisons: costs, real-time granularity, and data sovereignty for every major tool.

See why teams switch from Datadog, Prometheus, Grafana, and more.

> Browse all comparisons
Edge-Native Observability, Born Open Source
Per-second visibility, ML on every metric, and data that never leaves your infrastructure.
Founded in 2016
615+ contributors worldwide
Remote-first, engineering-driven
Open source first
> Read our story
Promises We Publish—and Prove
12 principles backed by open code, independent validation, and measurable outcomes.
Open source, peer-reviewed
Zero config, instant value
Data sovereignty by design
Aligned pricing, no surprises
> See all 12 principles
Edge-Native, AI-Ready, 100% Open
76k+ stars. Full ML, AI, and automation—GPLv3+, not premium add-ons.
76,000+ GitHub stars
GPLv3+ licensed forever
ML on every metric, included
Zero vendor lock-in
> Explore our open source
Build Real-Time Observability for the World
Remote-first team shipping per-second monitoring with ML on every metric.
Remote-first, fully distributed
Open source (76k+ stars)
Challenging technical problems
Your code on millions of systems
> See open roles
Meet the Team Behind Netdata
Conferences, meetups, and tradeshows where you can see Netdata in action and talk to the engineers who build it.
Live demos and deep dives
Book 1-on-1 meetings
Talks and panel sessions
Event recaps and photos
> See all events
Talk to a Netdata Human in <24 Hours
Sales, partnerships, press, or professional services—real engineers, fast answers.
Discuss your observability needs
Pricing and volume discounts
Partnership opportunities
Media and press inquiries
> Book a conversation
Your Data. Your Rules.
On-prem data, cloud control plane, transparent terms.
Trust & Scale
76,000+ GitHub Stars

One of the most popular open-source monitoring projects

SOC 2 Type 2 Certified

Enterprise-grade security and compliance

Data Sovereignty

Your metrics stay on your infrastructure

Validated
University of Amsterdam

"Most energy-efficient monitoring solution" — ICSOC 2023, peer-reviewed

ADASTEC (Autonomous Driving)

"Doesn't miss alerts—mission-critical trust for safety software"

Community Stats
615+ Contributors

Global community improving monitoring for everyone

1.5M+ Downloads/Day

Trusted by teams worldwide

GPLv3+ Licensed

Free forever, fully open source agent

Why Join?
Remote-First

Work from anywhere, async-friendly culture

Impact at Scale

Your work helps millions of systems

$ guides / bind-dns / bind-dns-clients-per-query ▌

Operations Guides

BIND clients-per-query and max-clients-per-query: duplicate recursion for popular names

When hundreds of clients query the same domain at the same instant, BIND does not send hundreds of identical recursive queries upstream. It sends one fetch and attaches the waiting clients to it. Two configuration knobs control how many waiters can attach before BIND starts dropping the overflow: clients-per-query (the soft limit, default 10) and max-clients-per-query (the hard ceiling, default 100).

How it works

BIND deduplicates concurrent identical queries at the fetch level. When the first client queries a name that requires recursion, BIND opens a single fetch to the upstream authoritative server. Subsequent clients querying the same name, type, and class while that fetch is in flight do not trigger additional upstream queries. They attach as waiters to the existing fetch. When the fetch completes, all attached waiters receive the answer.

Without deduplication, a cache miss on a popular name would generate one upstream query per waiting client. A single expiring TTL on a hot domain could produce a burst of hundreds of simultaneous upstream fetches for the same record, consuming recursive-client slots, file descriptors, and upstream bandwidth for redundant work.

The deduplication mechanism has three layers: the initial soft limit, the auto-tuning algorithm that adjusts the effective limit based on observed demand, and the hard ceiling that caps the auto-tuning.

flowchart TD
    A[Query for uncached name] --> B{Identical fetch
already in flight?} B -->|No, first querier| C[Send fetch upstream
Hold 1 recursive-client slot] B -->|Yes, duplicate| D{Attached waiters
below spillat?} D -->|Yes| E[Attach to existing fetch] D -->|No| F[Drop query
Counted in QryDropped] C --> G[Fetch completes] G --> H{Succeeded?} H -->|Yes| I[Answer all waiters
Raise spillat by 5] H -->|No| J[Fail all waiters
spillat unchanged]

The soft limit: clients-per-query (default 10)

When a new query arrives for a name that already has an in-flight fetch, BIND checks how many clients are already attached as waiters. If the waiter count is below the current effective limit, the new client attaches and no upstream query is sent. If the waiter count is at or above the effective limit, the new query is dropped.

The hard ceiling: max-clients-per-query (default 100)

The absolute maximum number of clients that can wait on a single in-flight fetch, regardless of auto-tuning. The auto-tuned limit never exceeds this value.

The auto-tuning algorithm

BIND adjusts the effective limit between clients-per-query and max-clients-per-query based on observed demand. The resolver source tracks the effective limit internally as spillat; it is not a configuration option.

The algorithm works as follows:

  1. When a fetch completes successfully and clients were attached as waiters, BIND raises spillat by 5. The next time a popular name triggers a burst, more clients can attach before being dropped.
  2. After 20 minutes with no drops due to the clients-per-query quota, BIND lowers spillat by 1, back toward the configured soft limit.
  3. spillat never exceeds max-clients-per-query.
  4. If the fetch does not complete successfully (upstream authoritative server unreachable or timed out), spillat is not raised. This prevents BIND from increasing the number of clients waiting on a name that is slow or impossible to resolve, since those waiters would all be held until the fetch times out or fails.

Disabling the limit

Setting either option to 0 removes the corresponding bound. With clients-per-query 0, there is no soft limit. With max-clients-per-query 0, there is no hard ceiling. The only remaining constraint is the global recursive-clients limit (default 1000, soft quota at 90%, meaning 900).

Where drops surface in statistics

Queries dropped because the clients-per-query quota is exceeded are counted in the QryDropped counter in NSStats. QryDropped is a superset that includes drops from multiple fetch-limit mechanisms:

Drop sourceWhat it limits
clients-per-query / max-clients-per-queryWaiters per single in-flight fetch
fetches-per-zoneSimultaneous fetches per domain
fetches-per-serverSimultaneous fetches per upstream server
rate-limit (RRL)Response rate limiting

RateDropped counts only RRL-related drops. To isolate clients-per-query drops, check ClientQuota, which counts queries spilled for exceeding the clients-per-query quota. Also confirm RateDropped is not increasing proportionally.

BIND has a spill logging category for queries terminated by fetch-limit quotas. Enabling this logging during investigation can confirm whether drops are specifically from clients-per-query or from another fetch limit.

Where it shows up in production

Thundering herd on popular names

The most common scenario is TTL expiry on a hot domain. When a popular name’s cache entry expires, the next query triggers a fetch. If dozens or hundreds of clients query the same name simultaneously, they all arrive while the fetch is in flight. With the default clients-per-query of 10, once the waiter count reaches the current effective limit, additional clients are dropped if spillat has not yet been raised.

Common triggers:

  • CDN edge records with short TTLs queried by many application instances
  • Service discovery queries where many containers start simultaneously
  • DNS-based load balancer records with low TTLs

Slow upstream amplification

When an upstream authoritative server is slow to respond, each in-flight fetch holds its waiters for the full fetch duration. If the upstream is consistently slow, spillat does not auto-tune upward because the fetch does not complete successfully. The deduplication ceiling stays at the configured soft limit, causing more drops during upstream degradation. This is correct behavior: BIND should not pile up hundreds of clients waiting on an unreachable upstream.

Misconfiguration: clients-per-query greater than max-clients-per-query

Prior to BIND 9.20.8, if you set clients-per-query higher than max-clients-per-query without also raising the ceiling, BIND accepted the configuration silently. The auto-tuning algorithm could not function because its minimum exceeded its maximum. The soft limit would never adjust upward.

BIND 9.20.8 fixed this via GL #5224: if max-clients-per-query is set lower than clients-per-query, the value is silently adjusted upward to match clients-per-query. If you are running BIND 9.20.7 or earlier and have raised clients-per-query above 100 without also raising max-clients-per-query, your auto-tuning is effectively stuck at the floor.

Tradeoffs

Raising the limits

Increasing clients-per-query and max-clients-per-query allows more clients to benefit from deduplication during popular-name bursts. This reduces upstream query volume and improves response latency for clients that would otherwise be dropped and forced to retry.

The cost is resource consumption. Each client waiting on a fetch holds state in BIND’s memory. Each in-flight fetch holds a slot in the recursive-clients table (default 1000). More waiters per fetch means more state tied to each fetch; BIND does not document a fixed per-waiter file-descriptor cost.

If you raise max-clients-per-query significantly, also verify:

  • Your recursive-clients limit is sufficient for the combined load of unique fetches plus their waiter pools.
  • Your file descriptor limit has headroom. The BIND files option is deprecated in BIND 9.18 and removed in BIND 9.20; FD limits are OS-controlled via ulimit or systemd LimitNOFILE.
  • The upstream authoritative servers can handle the burst of queries that deduplication was previously suppressing. Raising limits means fewer dropped clients locally but more upstream load when the deduplicated queries finally fire.

Lowering the limits

Lowering clients-per-query below the default makes BIND more aggressive about dropping duplicate queries. This can protect the resolver during upstream degradation by limiting how many clients are held waiting on slow fetches. The tradeoff is more client-visible drops during normal popular-name bursts, increasing perceived latency for clients that must retry.

When to leave the defaults

For most recursive resolvers, the defaults (10 and 100) are reasonable. The auto-tuning algorithm adapts the effective limit based on actual demand, so manual tuning is only needed when:

  • You observe sustained QryDropped increases correlated with popular-name bursts.
  • You have confirmed via the spill logging category or process of elimination that drops are from clients-per-query, not RRL or fetches-per-zone.
  • Your workload has a specific pattern the defaults do not handle well, such as a very high concentration of queries for a small set of names with short TTLs.

Signals to watch

SignalWhy it mattersWarning sign
QryDropped (NSStats)Superset counter for all fetch-limit drops including clients-per-querySustained increase correlated with popular-name cache-miss bursts
RateDropped (NSStats)RRL-specific drops; helps isolate whether QryDropped increase is RRL or clients-per-queryFlat RateDropped while QryDropped rises indicates fetch limits are the cause
RecursClients (NSStats)Global recursive client pressure; each fetch holds a slotRising toward limit alongside popular-name bursts
NumFetch (per-view resolver stats)Per-view active fetch countHigh NumFetch with moderate RecursClients suggests many waiters per fetch
File descriptor usageEach waiter holds state; FD headroom needed for upstream socketsFD usage approaching limit after raising clients-per-query
Cache hit ratioFalling hit ratio increases cache misses, increasing dedup pressureDeclining hit ratio alongside rising QryDropped

How Netdata helps

  • Per-second QryDropped and RateDropped collection lets you pinpoint the exact moment drops begin and correlate them with cache-miss bursts, TTL expirations, or upstream RTT shifts. Thundering-herd events can be brief enough that minute-level polling misses them.
  • RecursClients as a live gauge shows whether deduplication drops are happening because the global recursive-client table is also saturated, or whether the issue is isolated to the per-query dedup limit.
  • Per-view NumFetch reveals whether deduplication pressure is concentrated in one view, common in split-horizon deployments where internal clients hammer service discovery names.
  • Cache hit ratio trends show whether declining cache effectiveness is driving more cache misses and therefore more deduplication events.
  • Upstream RTT distribution (QryRTT* buckets) shows whether slow upstream responses are preventing spillat from auto-tuning upward, keeping the effective dedup limit pinned at the floor.
  • File descriptor monitoring confirms whether raising max-clients-per-query is safe given current FD headroom on the host.