The only agent that thinks for itself

Autonomous Monitoring with self-learning AI built-in, operating independently across your entire stack.

Unlimited Metrics & Logs
Machine learning & MCP
5% CPU, 150MB RAM
3GB disk, >1 year retention
800+ integrations, zero config
Dashboards, alerts out of the box
> Discover Netdata Agents

Centralized metrics streaming and storage

Aggregate metrics from multiple agents into centralized Parent nodes for unified monitoring across your infrastructure.

Stream from unlimited agents
Long-term data retention
High availability clustering
Data replication & backup
Scalable architecture
Enterprise-grade security
> Learn about Parents

Fully managed cloud platform

Access your monitoring data from anywhere with our SaaS platform. No infrastructure to manage, automatic updates, and global availability.

Zero infrastructure management
99.9% uptime SLA
Global data centers
Automatic updates & patches
Enterprise SSO & RBAC
SOC2 & ISO certified
> Explore Netdata Cloud

Deploy Netdata Cloud in your infrastructure

Run the full Netdata Cloud platform on-premises for complete data sovereignty and compliance with your security policies.

Complete data sovereignty
Air-gapped deployment
Custom compliance controls
Private network integration
Dedicated support team
Kubernetes & Docker support
> Learn about Cloud On-Premises

Powerful, intuitive monitoring interface

Modern, responsive UI built for real-time troubleshooting with customizable dashboards and advanced visualization capabilities.

Real-time chart updates
Customizable dashboards
Dark & light themes
Advanced filtering & search
Responsive on all devices
Collaboration features
> Explore Netdata UI

Monitor on the go

Native iOS and Android apps bring full monitoring capabilities to your mobile device with real-time alerts and notifications.

iOS & Android apps
Push notifications
Touch-optimized interface
Offline data access
Biometric authentication
Widget support
> Download apps

The future of infrastructure observability

See our strategic direction across AI-native observability, full-stack signals, operational intelligence, and enterprise platform maturity.

AI-native observability
Full-stack signal coverage
Operational intelligence
Enterprise platform maturity
Agent releases every 6 weeks
Cloud continuous delivery
> Explore Product Roadmap

Best energy efficiency

True real-time per-second

100% automated zero config

Centralized observability

Multi-year retention

High availability built-in

Zero maintenance

Always up-to-date

Enterprise security

Complete data control

Air-gap ready

Compliance certified

Millisecond responsiveness

Infinite zoom & pan

Works on any device

Native performance

Instant alerts

Monitor anywhere

AI-native observability

Continuous delivery

Open source foundation

80% Faster Incident Resolution

AI-powered troubleshooting from detection, to root cause and blast radius identification, to reporting.

True Real-Time and Simple, even at Scale

Linearly and infinitely scalable full-stack observability, that can be deployed even mid-crisis.

90% Cost Reduction, Full Fidelity

Instead of centralizing the data, Netdata distributes the code, eliminating pipelines and complexity.

See and Map Your Entire Network

Live topology, flow analytics, and SNMP device and trap monitoring — unified with your full-stack observability.

Control Without Surrender

SOC 2 Type 2 certified with every metric kept on your infrastructure.

Integrations

800+ collectors and notification channels, auto-discovered and ready out of the box.

800+ data collectors
Auto-discovery & zero config
Cloud, infra, app protocols
Notifications out of the box
> Explore integrations
Real Results
46% Cost Reduction

Reduced monitoring costs by 46% while cutting staff overhead by 67%.

— Leonardo Antunez, Codyas

Zero Pipeline

No data shipping. No central storage costs. Query at the edge.

From Our Users
"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

No Query Language

Point-and-click troubleshooting. No PromQL, no LogQL, no learning curve.

Enterprise Ready
67% Less Staff, 46% Cost Cut

Enterprise efficiency without enterprise complexity—real ROI from day one.

— Leonardo Antunez, Codyas

SOC 2 Type 2 Certified

Zero data egress. Only metadata reaches the cloud. Your metrics stay on your infrastructure.

Full Coverage
800+ Collectors

Auto-discovered and configured. No manual setup required.

Any Notification Channel

Slack, PagerDuty, Teams, email, webhooks—all built-in.

Built for the People Who Get Paged

Because 3am alerts deserve instant answers, not hour-long hunts.

Every Industry Has Rules. We Master Them.

See how healthcare, finance, and government teams cut monitoring costs 90% while staying audit-ready.

Monitor Any Technology. Configure Nothing.

Install the agent. It already knows your stack.
From Our Users
"A Rare Unicorn"

Netdata gives more than you invest in it. A rare unicorn that obeys the Pareto rule.

— Eduard Porquet Mateu, TMB Barcelona

99% Downtime Reduction

Reduced website downtime by 99% and cloud bill by 30% using Netdata alerts.

— Falkland Islands Government

Real Savings
30% Cloud Cost Reduction

Optimized resource allocation based on Netdata alerts cut cloud spending by 30%.

— Falkland Islands Government

46% Cost Cut

Reduced monitoring staff by 67% while cutting operational costs by 46%.

— Codyas

Real Coverage
"Plugin for Everything"

Netdata has agent capacity or a plugin for everything, including Windows and Kubernetes.

— Eduard Porquet Mateu, TMB Barcelona

"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

Real Speed
Troubleshooting in 30 Seconds

From 2-3 minutes to 30 seconds—instant visibility into any node issue.

— Matthew Artist, Nodecraft

20% Downtime Reduction

20% less downtime and 40% budget optimization from out-of-the-box monitoring.

— Simon Beginn, LANCOM Systems

Pay per Node. Unlimited Everything Else.

One price per node. Unlimited metrics, logs, users, and retention. No per-GB surprises.

Free tier—forever
No metric limits or caps
Retention you control
Cancel anytime
> See pricing plans

What's Your Monitoring Really Costing You?

Most teams overpay by 40-60%. Let's find out why.

Expose hidden metric charges
Calculate tool consolidation
Customers report 30-67% savings
Results in under 60 seconds
> See what you're really paying

Your Infrastructure Is Unique. Let's Talk.

Because monitoring 10 nodes is different from monitoring 10,000.

On-prem & air-gapped deployment
Volume pricing & agreements
Architecture review for your scale
Compliance & security support
> Start a conversation

Monitoring That Sells Itself

Deploy in minutes. Impress clients in hours. Earn recurring revenue for years.

30-second live demos close deals
Zero config = zero support burden
Competitive margins & deal protection
Response in 48 hours
> Apply to partner

Per-Second Metrics at Homelab Prices

Same engine, same dashboards, same ML. Just priced for tinkerers.

Community: Free forever · 5 nodes · non-commercial
Homelab: $90/yr · unlimited nodes · fair usage
> Get the Homelab Plan

$1,000 Per Referral. Unlimited Referrals.

Your colleagues get 10% off. You get 10% commission. Everyone wins.

10% of subscriptions, up to $1,000 each
Track earnings inside Netdata Cloud
PayPal/Venmo payouts in 3-4 weeks
No caps, no complexity
> Get your referral link
Cost Proof
40% Budget Optimization

"Netdata's significant positive impact" — LANCOM Systems

Calculate Your Savings

Compare vs Datadog, Grafana, Dynatrace

Savings Proof
46% Cost Reduction

"Cut costs by 46%, staff by 67%" — Codyas

30% Cloud Bill Savings

"Reduced cloud bill by 30%" — Falkland Islands Gov

Enterprise Proof
"Better Than Combined Alternatives"

"Better observability with Netdata than combining other tools." — TMB Barcelona

Real Engineers, <24h Response

DPA, SLAs, on-prem, volume pricing

Why Partners Win
Demo Live Infrastructure

One command, 30 seconds, real data—no sandbox needed

Zero Tickets, High Margins

Auto-config + per-node pricing = predictable profit

Homelab Ready
Free Video Course

8-episode Netdata tutorial by LearnLinux.tv

76k+ GitHub Stars

3rd most starred monitoring project

Worth Recommending
Product That Delivers

Customers report 40-67% cost cuts, 99% downtime reduction

Zero Risk to Your Rep

Free tier lets them try before they buy

AI Support Assistant, Available 24/7

Nedi has access to all official documentation, source code, and resources. Ask any question about Netdata—responds in your language.

Deployment & configuration
Troubleshooting & sizing
Alerts & notifications
Evidence-based answers
> Ask Nedi now

Never Fight Fires Alone

Docs, community, and expert help—pick your path to resolution.

Learn.netdata.cloud docs
Discord, Forums, GitHub
Premium support available
> Get answers now

60 Seconds to First Dashboard

One command to install. Zero config. 850+ integrations documented.

Linux, Windows, K8s, Docker
Auto-discovers your stack
> Read our documentation

76,000+ Engineers Strong

615+ contributors. 1.5M daily downloads. One mission: simplify observability.

Per-Second. 90% Cheaper. Data Stays Home.

Side-by-side comparisons: costs, real-time granularity, and data sovereignty for every major tool.

See why teams switch from Datadog, Prometheus, Grafana, and more.

> Browse all comparisons
Edge-Native Observability, Born Open Source
Per-second visibility, ML on every metric, and data that never leaves your infrastructure.
Founded in 2016
615+ contributors worldwide
Remote-first, engineering-driven
Open source first
> Read our story
Promises We Publish—and Prove
12 principles backed by open code, independent validation, and measurable outcomes.
Open source, peer-reviewed
Zero config, instant value
Data sovereignty by design
Aligned pricing, no surprises
> See all 12 principles
Edge-Native, AI-Ready, 100% Open
76k+ stars. Full ML, AI, and automation—GPLv3+, not premium add-ons.
76,000+ GitHub stars
GPLv3+ licensed forever
ML on every metric, included
Zero vendor lock-in
> Explore our open source
Build Real-Time Observability for the World
Remote-first team shipping per-second monitoring with ML on every metric.
Remote-first, fully distributed
Open source (76k+ stars)
Challenging technical problems
Your code on millions of systems
> See open roles
Meet the Team Behind Netdata
Conferences, meetups, and tradeshows where you can see Netdata in action and talk to the engineers who build it.
Live demos and deep dives
Book 1-on-1 meetings
Talks and panel sessions
Event recaps and photos
> See all events
Talk to a Netdata Human in <24 Hours
Sales, partnerships, press, or professional services—real engineers, fast answers.
Discuss your observability needs
Pricing and volume discounts
Partnership opportunities
Media and press inquiries
> Book a conversation
Your Data. Your Rules.
On-prem data, cloud control plane, transparent terms.
Trust & Scale
76,000+ GitHub Stars

One of the most popular open-source monitoring projects

SOC 2 Type 2 Certified

Enterprise-grade security and compliance

Data Sovereignty

Your metrics stay on your infrastructure

Validated
University of Amsterdam

"Most energy-efficient monitoring solution" — ICSOC 2023, peer-reviewed

ADASTEC (Autonomous Driving)

"Doesn't miss alerts—mission-critical trust for safety software"

Community Stats
615+ Contributors

Global community improving monitoring for everyone

1.5M+ Downloads/Day

Trusted by teams worldwide

GPLv3+ Licensed

Free forever, fully open source agent

Why Join?
Remote-First

Work from anywhere, async-friendly culture

Impact at Scale

Your work helps millions of systems

$ guides / bind-dns / bind-dns-numfetch-per-view ▌

Operations Guides

BIND resolver NumFetch per view: per-view recursive pressure in split-horizon setups

RecursClients is a global NSStats counter that reports the total number of recursive clients awaiting resolution across all views. It does not break down which view is consuming the pool. When RecursClients climbs to 847 out of 1000, you know the pool is stressed but not whether the pressure is evenly distributed or whether one view is responsible for most of it.

NumFetch per view fills that gap. It is a view-scoped gauge of active outbound fetches to upstream nameservers, maintained independently for each view in the resolver statistics. Unlike cumulative counters that increment over time, NumFetch is an instantaneous gauge: it goes up when a fetch starts and down when the fetch completes, times out, or is aborted. A sustained upward trend in one view’s NumFetch while other views remain flat localizes a view-specific resolution problem that the global counter masks.

Each active fetch holds resources (a recursive-client slot, a file descriptor, memory for resolver state) until it completes or times out. When upstream nameservers respond quickly, fetches complete rapidly and NumFetch stays low. When upstream is slow or unreachable, fetches pile up and NumFetch climbs.

There is no per-view recursive-clients limit. The option is global, set only in the options block, and caps the sum across all views. One view’s runaway recursive load can exhaust the shared pool and starve other views.

How it works

Every DNS query arrives at BIND through view selection first. The match-clients and match-destinations directives in each view block determine which view processes the query. Once a view accepts a query and determines it needs recursion (cache miss, recursion enabled), the view’s resolver initiates one or more outbound fetches. NumFetch counts these active fetches per view.

All per-view fetches draw from the same recursive-clients pool. BIND maintains a single global table of recursive client slots, bounded by recursive-clients (default 1000). When the table reaches the soft quota (90%, or 900 by default), BIND logs a rate-limited warning and reclaims the oldest slot; the incoming recursive query is then served. At the hard limit (1000), BIND logs a rate-limited warning, reclaims the oldest slot, and the new recursive query fails with SERVFAIL.

This shared-pool architecture means one view’s slow upstream dependency can cascade into a global failure:

flowchart TD
    CI["internal clients"] --> VI["View: internal
NumFetch: 45"] CE["external clients"] --> VE["View: external
NumFetch: 780"] CG["guest clients"] --> VG["View: guest
NumFetch: 22"] VI -->|"recursion"| RC["Shared recursive-clients
847 of 1000 slots"] VE -->|"recursion"| RC VG -->|"recursion"| RC RC --> SQ["soft quota: 900"] RC --> HL["hard limit: 1000
SERVFAIL for all"]

If the external view’s forwarder becomes unreachable, every query the external view sends upstream occupies a slot for the resolver-query-timeout duration (default 10 seconds). As those slots accumulate, fewer remain for the internal and guest views. Eventually all views suffer SERVFAIL even though only one view has an upstream problem. If external-view NumFetch is climbing while internal-view NumFetch stays flat, you have localized the problem before RecursClients reaches the soft quota.

Collecting NumFetch per view

The statistics channel exposes NumFetch under each view’s resolver stats. The port depends on your statistics-channels configuration (commonly 8053 or 8653):

# Dump NumFetch for every view from the JSON statistics channel
curl -s http://localhost:8653/json/v1/server | \
  python3 -c "import sys,json; d=json.load(sys.stdin); \
  [print(f'{v}: NumFetch={vd.get(\"resolver\",{}).get(\"stats\",{}).get(\"NumFetch\",\"N/A\")}') \
  for v,vd in d.get('views',{}).items()]"

If the statistics channel is not enabled, rndc stats writes per-view resolver statistics (including NumFetch) to the configured statistics file. The path varies by distribution; check your statistics-file setting:

rndc stats
grep NumFetch /var/named/data/named_stats.txt

During an active incident, rndc recursing shows the queries currently recursing and the domains with active iterative fetches. This is the next level of detail after NumFetch identifies which view is under pressure:

rndc recursing

Version notes

Verify counter names against the statistics output for your BIND release. Zero-valued counters may be omitted from non-verbose statistics-channel output, so a view with NumFetch at 0 may not appear unless verbose mode is enabled.

BIND 9.18.1 fixed a GL #3147 bug in which the global RecursClients counter could be miscalculated and drop below zero in certain resolution scenarios. NumFetch was not identified as affected. If you are running BIND 9.18.0 and see impossible RecursClients values, upgrade to 9.18.1 or later.

When it matters

NumFetch per view matters in any deployment where multiple views share a single named process and a single recursive-clients pool:

  • Split-horizon resolvers with separate internal and external views. The internal view may forward to corporate DNS while the external view resolves against public nameservers. A failure in either forwarding path shows up as elevated NumFetch in only that view.
  • Multi-tenant DNS where different client classes receive different views. One tenant’s misconfigured application generating excessive unique queries drives up that view’s NumFetch without affecting others.
  • Guest or DMZ networks resolving through a different upstream path than production networks. A guest-network upstream degradation is isolated in the guest view’s NumFetch.
  • Mixed-role instances serving both authoritative zones and recursive resolution through different views. Recursive pressure in the resolving view is distinguishable from authoritative query load.

Use NumFetch per view when:

  • RecursClients is trending toward 50% of the limit and you need early warning about which view is the heaviest consumer.
  • One view’s clients report intermittent DNS failures while others are unaffected.
  • A forwarder change affects one view’s upstream path and you need to confirm the impact.
  • You are planning capacity and need to know which view drives recursive load growth.

The global counter suffices when you run a single-view resolver, or when all views share the same upstream path and a homogeneous client population. If you are monitoring specifically for the recursive-clients circuit breaker, RecursClients as a percentage of the limit is the right signal for that.

Important limitation: there is no per-view recursive-clients cap. You cannot configure BIND to limit one view to 300 recursive clients and another to 700. The global limit applies to the sum. If one view’s traffic is consuming a disproportionate share, the response is operational (investigate and mitigate the upstream cause in that view) rather than configurational (set a per-view quota).

Signals to watch

NumFetch per view is most useful when correlated with other per-view resolver signals.

SignalWhy it mattersWarning sign
NumFetch per viewLocalizes recursive pressure to a specific viewOne view trending up while others stay flat
RecursClients (global)Shows total recursive pool utilization against the limitApproaching 900 (soft quota) or 1000 (hard limit)
Per-view cache hit ratioFalling hit ratio drives more fetches in that viewDecline in the same view where NumFetch is rising
Per-view RTT distributionSlow upstream increases fetch duration, holding slots longerShift toward higher RTT buckets (QryRTT1600+) in one view
Per-view QueryTimeoutTimeouts hold slots for the full resolver-query-timeoutRate exceeding 5% of outbound queries in one view
Per-view CacheMissesHigh miss rate means more outbound fetchesSpike concentrated in one view, possibly from random subdomain queries

The diagnostic chain: NumFetch identifies the view under pressure. Cache hit ratio in that view tells you whether the pressure is from cache misses (more queries need recursion) or slow upstream (the same fetches take longer). RTT distribution confirms upstream latency. QueryTimeout tells you whether upstream is not responding at all.

How Netdata helps

  • Collects NumFetch per view at per-second resolution from the BIND statistics channel, without manual polling or delta computation.
  • Correlates per-view NumFetch with the global RecursClients gauge in the same dashboard, showing what fraction of the shared pool each view is consuming at any moment.
  • Surfaces per-view cache hit ratio, RTT distribution, and QueryTimeout alongside NumFetch, making the cause-effect chain visible without switching tools.
  • Applies ML anomaly detection to per-view NumFetch trends, flagging a sustained upward drift in one view before the global RecursClients crosses a threshold.