The only agent that thinks for itself

Autonomous Monitoring with self-learning AI built-in, operating independently across your entire stack.

Unlimited Metrics & Logs
Machine learning & MCP
5% CPU, 150MB RAM
3GB disk, >1 year retention
800+ integrations, zero config
Dashboards, alerts out of the box
> Discover Netdata Agents

Centralized metrics streaming and storage

Aggregate metrics from multiple agents into centralized Parent nodes for unified monitoring across your infrastructure.

Stream from unlimited agents
Long-term data retention
High availability clustering
Data replication & backup
Scalable architecture
Enterprise-grade security
> Learn about Parents

Fully managed cloud platform

Access your monitoring data from anywhere with our SaaS platform. No infrastructure to manage, automatic updates, and global availability.

Zero infrastructure management
99.9% uptime SLA
Global data centers
Automatic updates & patches
Enterprise SSO & RBAC
SOC2 & ISO certified
> Explore Netdata Cloud

Deploy Netdata Cloud in your infrastructure

Run the full Netdata Cloud platform on-premises for complete data sovereignty and compliance with your security policies.

Complete data sovereignty
Air-gapped deployment
Custom compliance controls
Private network integration
Dedicated support team
Kubernetes & Docker support
> Learn about Cloud On-Premises

Powerful, intuitive monitoring interface

Modern, responsive UI built for real-time troubleshooting with customizable dashboards and advanced visualization capabilities.

Real-time chart updates
Customizable dashboards
Dark & light themes
Advanced filtering & search
Responsive on all devices
Collaboration features
> Explore Netdata UI

Monitor on the go

Native iOS and Android apps bring full monitoring capabilities to your mobile device with real-time alerts and notifications.

iOS & Android apps
Push notifications
Touch-optimized interface
Offline data access
Biometric authentication
Widget support
> Download apps

The future of infrastructure observability

See our strategic direction across AI-native observability, full-stack signals, operational intelligence, and enterprise platform maturity.

AI-native observability
Full-stack signal coverage
Operational intelligence
Enterprise platform maturity
Agent releases every 6 weeks
Cloud continuous delivery
> Explore Product Roadmap

Best energy efficiency

True real-time per-second

100% automated zero config

Centralized observability

Multi-year retention

High availability built-in

Zero maintenance

Always up-to-date

Enterprise security

Complete data control

Air-gap ready

Compliance certified

Millisecond responsiveness

Infinite zoom & pan

Works on any device

Native performance

Instant alerts

Monitor anywhere

AI-native observability

Continuous delivery

Open source foundation

80% Faster Incident Resolution

AI-powered troubleshooting from detection, to root cause and blast radius identification, to reporting.

True Real-Time and Simple, even at Scale

Linearly and infinitely scalable full-stack observability, that can be deployed even mid-crisis.

90% Cost Reduction, Full Fidelity

Instead of centralizing the data, Netdata distributes the code, eliminating pipelines and complexity.

See and Map Your Entire Network

Live topology, flow analytics, and SNMP device and trap monitoring — unified with your full-stack observability.

Control Without Surrender

SOC 2 Type 2 certified with every metric kept on your infrastructure.

Integrations

800+ collectors and notification channels, auto-discovered and ready out of the box.

800+ data collectors
Auto-discovery & zero config
Cloud, infra, app protocols
Notifications out of the box
> Explore integrations
Real Results
46% Cost Reduction

Reduced monitoring costs by 46% while cutting staff overhead by 67%.

— Leonardo Antunez, Codyas

Zero Pipeline

No data shipping. No central storage costs. Query at the edge.

From Our Users
"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

No Query Language

Point-and-click troubleshooting. No PromQL, no LogQL, no learning curve.

Enterprise Ready
67% Less Staff, 46% Cost Cut

Enterprise efficiency without enterprise complexity—real ROI from day one.

— Leonardo Antunez, Codyas

SOC 2 Type 2 Certified

Zero data egress. Only metadata reaches the cloud. Your metrics stay on your infrastructure.

Full Coverage
800+ Collectors

Auto-discovered and configured. No manual setup required.

Any Notification Channel

Slack, PagerDuty, Teams, email, webhooks—all built-in.

Built for the People Who Get Paged

Because 3am alerts deserve instant answers, not hour-long hunts.

Every Industry Has Rules. We Master Them.

See how healthcare, finance, and government teams cut monitoring costs 90% while staying audit-ready.

Monitor Any Technology. Configure Nothing.

Install the agent. It already knows your stack.
From Our Users
"A Rare Unicorn"

Netdata gives more than you invest in it. A rare unicorn that obeys the Pareto rule.

— Eduard Porquet Mateu, TMB Barcelona

99% Downtime Reduction

Reduced website downtime by 99% and cloud bill by 30% using Netdata alerts.

— Falkland Islands Government

Real Savings
30% Cloud Cost Reduction

Optimized resource allocation based on Netdata alerts cut cloud spending by 30%.

— Falkland Islands Government

46% Cost Cut

Reduced monitoring staff by 67% while cutting operational costs by 46%.

— Codyas

Real Coverage
"Plugin for Everything"

Netdata has agent capacity or a plugin for everything, including Windows and Kubernetes.

— Eduard Porquet Mateu, TMB Barcelona

"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

Real Speed
Troubleshooting in 30 Seconds

From 2-3 minutes to 30 seconds—instant visibility into any node issue.

— Matthew Artist, Nodecraft

20% Downtime Reduction

20% less downtime and 40% budget optimization from out-of-the-box monitoring.

— Simon Beginn, LANCOM Systems

Pay per Node. Unlimited Everything Else.

One price per node. Unlimited metrics, logs, users, and retention. No per-GB surprises.

Free tier—forever
No metric limits or caps
Retention you control
Cancel anytime
> See pricing plans

What's Your Monitoring Really Costing You?

Most teams overpay by 40-60%. Let's find out why.

Expose hidden metric charges
Calculate tool consolidation
Customers report 30-67% savings
Results in under 60 seconds
> See what you're really paying

Your Infrastructure Is Unique. Let's Talk.

Because monitoring 10 nodes is different from monitoring 10,000.

On-prem & air-gapped deployment
Volume pricing & agreements
Architecture review for your scale
Compliance & security support
> Start a conversation

Monitoring That Sells Itself

Deploy in minutes. Impress clients in hours. Earn recurring revenue for years.

30-second live demos close deals
Zero config = zero support burden
Competitive margins & deal protection
Response in 48 hours
> Apply to partner

Per-Second Metrics at Homelab Prices

Same engine, same dashboards, same ML. Just priced for tinkerers.

Community: Free forever · 5 nodes · non-commercial
Homelab: $90/yr · unlimited nodes · fair usage
> Get the Homelab Plan

$1,000 Per Referral. Unlimited Referrals.

Your colleagues get 10% off. You get 10% commission. Everyone wins.

10% of subscriptions, up to $1,000 each
Track earnings inside Netdata Cloud
PayPal/Venmo payouts in 3-4 weeks
No caps, no complexity
> Get your referral link
Cost Proof
40% Budget Optimization

"Netdata's significant positive impact" — LANCOM Systems

Calculate Your Savings

Compare vs Datadog, Grafana, Dynatrace

Savings Proof
46% Cost Reduction

"Cut costs by 46%, staff by 67%" — Codyas

30% Cloud Bill Savings

"Reduced cloud bill by 30%" — Falkland Islands Gov

Enterprise Proof
"Better Than Combined Alternatives"

"Better observability with Netdata than combining other tools." — TMB Barcelona

Real Engineers, <24h Response

DPA, SLAs, on-prem, volume pricing

Why Partners Win
Demo Live Infrastructure

One command, 30 seconds, real data—no sandbox needed

Zero Tickets, High Margins

Auto-config + per-node pricing = predictable profit

Homelab Ready
Free Video Course

8-episode Netdata tutorial by LearnLinux.tv

76k+ GitHub Stars

3rd most starred monitoring project

Worth Recommending
Product That Delivers

Customers report 40-67% cost cuts, 99% downtime reduction

Zero Risk to Your Rep

Free tier lets them try before they buy

AI Support Assistant, Available 24/7

Nedi has access to all official documentation, source code, and resources. Ask any question about Netdata—responds in your language.

Deployment & configuration
Troubleshooting & sizing
Alerts & notifications
Evidence-based answers
> Ask Nedi now

Never Fight Fires Alone

Docs, community, and expert help—pick your path to resolution.

Learn.netdata.cloud docs
Discord, Forums, GitHub
Premium support available
> Get answers now

60 Seconds to First Dashboard

One command to install. Zero config. 850+ integrations documented.

Linux, Windows, K8s, Docker
Auto-discovers your stack
> Read our documentation

76,000+ Engineers Strong

615+ contributors. 1.5M daily downloads. One mission: simplify observability.

Per-Second. 90% Cheaper. Data Stays Home.

Side-by-side comparisons: costs, real-time granularity, and data sovereignty for every major tool.

See why teams switch from Datadog, Prometheus, Grafana, and more.

> Browse all comparisons
Edge-Native Observability, Born Open Source
Per-second visibility, ML on every metric, and data that never leaves your infrastructure.
Founded in 2016
615+ contributors worldwide
Remote-first, engineering-driven
Open source first
> Read our story
Promises We Publish—and Prove
12 principles backed by open code, independent validation, and measurable outcomes.
Open source, peer-reviewed
Zero config, instant value
Data sovereignty by design
Aligned pricing, no surprises
> See all 12 principles
Edge-Native, AI-Ready, 100% Open
76k+ stars. Full ML, AI, and automation—GPLv3+, not premium add-ons.
76,000+ GitHub stars
GPLv3+ licensed forever
ML on every metric, included
Zero vendor lock-in
> Explore our open source
Build Real-Time Observability for the World
Remote-first team shipping per-second monitoring with ML on every metric.
Remote-first, fully distributed
Open source (76k+ stars)
Challenging technical problems
Your code on millions of systems
> See open roles
Meet the Team Behind Netdata
Conferences, meetups, and tradeshows where you can see Netdata in action and talk to the engineers who build it.
Live demos and deep dives
Book 1-on-1 meetings
Talks and panel sessions
Event recaps and photos
> See all events
Talk to a Netdata Human in <24 Hours
Sales, partnerships, press, or professional services—real engineers, fast answers.
Discuss your observability needs
Pricing and volume discounts
Partnership opportunities
Media and press inquiries
> Book a conversation
Your Data. Your Rules.
On-prem data, cloud control plane, transparent terms.
Trust & Scale
76,000+ GitHub Stars

One of the most popular open-source monitoring projects

SOC 2 Type 2 Certified

Enterprise-grade security and compliance

Data Sovereignty

Your metrics stay on your infrastructure

Validated
University of Amsterdam

"Most energy-efficient monitoring solution" — ICSOC 2023, peer-reviewed

ADASTEC (Autonomous Driving)

"Doesn't miss alerts—mission-critical trust for safety software"

Community Stats
615+ Contributors

Global community improving monitoring for everyone

1.5M+ Downloads/Day

Trusted by teams worldwide

GPLv3+ Licensed

Free forever, fully open source agent

Why Join?
Remote-First

Work from anywhere, async-friendly culture

Impact at Scale

Your work helps millions of systems

$ guides / bind-dns / bind-dns-too-many-open-files ▌

Operations Guides

BIND 'too many open files': file descriptor exhaustion and silently dropped queries

BIND logs “too many open files” or “socket: file descriptor exceeds limit” and starts dropping queries. UDP health checks may still pass. Zone transfers fail intermittently. The rndc control channel becomes sluggish or unresponsive. Clients see random timeouts that look like upstream nameserver problems. This is file descriptor exhaustion, and the symptoms masquerade as network or disk issues.

The default ulimit -n of 1024 on many Linux distributions is low for a production DNS server. BIND consumes file descriptors (FDs) for every listener socket, outbound recursive query socket, TCP client connection, zone file, journal file, log file, and the rndc control channel. Under moderate load, a recursive resolver can exhaust 1024 FDs in minutes.

The failure is cliff-edge. There is no graceful degradation. BIND is processing queries normally, then it cannot open new sockets and silently drops queries or fails transfers. UDP-only monitoring stays green because existing listener sockets remain open. The problem becomes visible only when you check FD usage against the limit or notice that TCP-dependent operations (zone transfers, large responses) are failing.

What this means

BIND’s FD consumption scales with:

  • Listener sockets: one UDP and one TCP socket per configured listen address, per IP version.
  • Outbound recursive sockets: each in-flight recursive query to an upstream nameserver holds at least one FD. The recursive-clients option (default 1000) caps concurrent recursive queries, but if the FD limit is lower than what recursive-clients allows, FDs exhaust first.
  • TCP client connections: each accepted TCP connection to port 53 holds an FD until closed. The tcp-clients option (default 150) limits concurrent TCP connections.
  • Zone and journal files: each loaded zone with dynamic updates or inline signing opens file handles for the zone file and its .jnl journal.
  • Log files: each configured logging channel that writes to a file holds an FD.
  • Control channel: the rndc listener on TCP port 953 holds an FD.

When the total exceeds Max open files, BIND cannot create new sockets or open new files. The kernel denies the socket(), accept(), or open() syscall. BIND logs one of several messages:

  • too many open files
  • socket: file descriptor exceeds limit (N/M) where N is current usage and M is the limit (BIND 9.16 and earlier)
  • accept: file descriptor exceeds limit (BIND 9.16 and earlier)

Not every failure produces a log line. Queries dropped during processing because BIND could not allocate a socket for the response may vanish silently. The gap between what BIND logs and what actually happens can be large.

flowchart TD
    A["Query arrives at named"] --> B{"FD available?"}
    B -- No --> C["Cannot open socket"]
    C --> D["Query silently dropped
No log entry"] C --> E["Transfer fails
Error logged"] C --> F["rndc degraded
or unresponsive"] B -- Yes --> G["Normal processing"] G --> H["FD held for
duration of operation"] H --> I["Operation completes
FD released"] H --> J{"More queries
than FDs released?"} J -- Yes --> K["FD count climbs
toward limit"] K --> B

Common causes

CauseWhat it looks likeFirst thing to check
Default ulimit too low (1024)Works under low load, fails under stress; FD usage near 100% of 1024 limitgrep "Max open files" /proc/$(pgrep -x named)/limits
systemd LimitNOFILE not set/etc/security/limits.conf was configured but named still has low limitsystemctl show named | grep LimitNOFILE
TCP connection accumulationTCP connections to port 53 rising; transfers or large responses failingss -tan '( sport = :53 )' | wc -l
Zone transfer burstMany secondaries transferring simultaneously; FD spike during refresh windowsBIND logs category xfer-in / xfer-out
Traffic growth exceeding capacityFD usage trending upward over weeks; daily peak approaching 50% of limitHistorical FD usage trend
FD leak (rare, version-specific)FD count grows monotonically without release; never stabilizes/proc/<pid>/fd count over time

Quick checks

Run these read-only commands to confirm or rule out FD exhaustion. The service name may be named or bind9 depending on distribution. If multiple named processes are running, pgrep -x named returns the first match; specify the PID explicitly instead.

# Check current FD count vs limit
PID=$(pgrep -x named)
CURRENT=$(ls /proc/$PID/fd 2>/dev/null | wc -l)
MAX=$(grep "Max open files" /proc/$PID/limits | awk '{print $4}')
echo "Using $CURRENT / $MAX file descriptors ($(( CURRENT * 100 / MAX ))%)"

# Check what named has open (socket targets, file paths)
ls -l /proc/$PID/fd 2>/dev/null | awk '{print $NF}' | sed 's/\[.*\]//' | sort | uniq -c | sort -rn | head -20

# Check for FD exhaustion errors in recent logs
journalctl -u named --since "1 hour ago" | grep -i "too many open files\|file descriptor exceeds limit\|not enough free resource"

# Check systemd's FD limit for named (overrides limits.conf for services)
systemctl show named | grep LimitNOFILE

# Check TCP connection count on port 53
ss -tan '( sport = :53 )' | tail -n +2 | awk '{print $1}' | sort | uniq -c

# Check QuerySockFail counter (socket errors on outbound queries)
# Requires a configured statistics-channels block in named.conf.
curl -s http://localhost:8653/json/v1/server | \
  python3 -c "import sys,json; d=json.load(sys.stdin); \
  [print(f'{v}: {k}: {s}') for v,vd in d.get('views',{}).items() \
  for k,s in sorted(vd.get('resolver',{}).get('stats',{}).items()) \
  if k == 'QuerySockFail']"

# Check recursive clients (each holds at least one FD)
curl -s http://localhost:8653/json/v1/server | \
  python3 -c "import sys,json; d=json.load(sys.stdin); \
  print('RecursClients:', d.get('nsstats',{}).get('RecursClients', 'N/A'))"

# Check rndc responsiveness (degrades under FD pressure)
timeout 5 rndc status >/dev/null 2>&1 && echo "rndc OK" || echo "rndc SLOW/FAIL"

# Functional TCP test (fails when FDs exhausted, UDP may still work)
dig +tcp +time=2 +tries=1 @127.0.0.1 example.com A

How to diagnose it

  1. Confirm FD usage is near the limit. Run the FD count and limit check above. Above 70% of Max open files is the warning zone. Above 90% is critical. At 100%, BIND is actively failing to open sockets.

  2. Identify what is consuming FDs. Examine /proc/$PID/fd to see the breakdown. Socket FDs (shown as socket:[NNN]) indicate listener, recursive, or TCP connections. Regular file paths indicate zone files, journals, or logs. A large count of socket FDs with many ESTABLISHED TCP connections points to TCP accumulation. A large count with high RecursClients points to recursive query load.

  3. Check whether systemd or limits.conf is the binding constraint. systemd’s LimitNOFILE takes precedence over /etc/security/limits.conf for services. If you set limits in limits.conf but did not configure the systemd unit, named still runs with the systemd default. Verify with systemctl show named | grep LimitNOFILE.

  4. Check the BIND version for the files option. The files option in named.conf was deprecated in BIND 9.18.10 (GL #3676) and removed in BIND 9.20.0, where using it is a configuration error. If named rejects the configuration on startup after an upgrade, remove the files directive and set FD limits at the OS level.

  5. Look for QuerySockFail in resolver statistics. This counter tracks failures opening query sockets for outbound recursive queries. A non-zero or increasing QuerySockFail rate is direct evidence that FD exhaustion is affecting recursive resolution.

  6. Test UDP vs TCP independently. If UDP queries work but TCP queries fail, FD exhaustion is likely. UDP listener sockets are long-lived and remain open. TCP connections require a new FD per accepted connection. When FDs are scarce, TCP fails first.

Metrics and signals to monitor

SignalWhy it mattersWarning sign
FD count as percentage of Max open filesDirect measure of resource exhaustionAbove 70% alert, above 90% page
TCP connection count on port 53TCP connections are FD-heavy; accumulation drives exhaustionSustained count near tcp-clients limit (default 150)
RecursClients as percentage of recursive-clientsEach in-flight recursive query holds FDsAbove 50% sustained with rising FD usage
QuerySockFail (per-view resolver stat)Direct evidence of socket creation failures from FD limitsAny non-zero rate
rndc status response timeControl channel uses FDs; degrades under pressureResponse time exceeding 5 seconds
QrySERVFAIL rateCollateral damage when BIND cannot process queriesSpike correlating with high FD usage
QryTCP vs QryUDP ratioElevated TCP share increases FD pressureTCP share above 5% of total queries

Fixes

Increase the OS file descriptor limit

The primary fix. On systemd-managed systems, create or edit the service override:

# Create systemd override for named
systemctl edit named

Add or modify:

[Service]
LimitNOFILE=65536

Then reload systemd and restart named:

systemctl daemon-reload
systemctl restart named

This restart interrupts DNS service briefly. Schedule accordingly.

For non-systemd systems, set limits in /etc/security/limits.conf:

named soft nofile 65536
named hard nofile 65536

The process must be restarted for the new limit to take effect. As root, you can raise limits on a running process with prlimit --pid <pid> --nofile=<new_soft>:<new_hard> as a temporary measure, but restart afterward so the systemd unit or limits.conf remains the source of truth.

Production BIND servers should run with at least 65536 FDs. High-traffic recursive resolvers or authoritative servers with many zones may need 1048576 or more.

Remove the deprecated files option from named.conf

The files option in named.conf was deprecated in BIND 9.18.10 and is a configuration error as of BIND 9.20.0 (GL #3676). Remove any files directive from named.conf and rely entirely on OS-level limits.

Reduce FD consumers if the limit cannot be raised

If you cannot raise the FD limit (container constraints, shared host), reduce FD consumption:

  • Lower tcp-clients (default 150). Each concurrent TCP connection holds an FD. Reducing this limits TCP FD consumption but may cause legitimate TCP queries to be refused.
  • Lower recursive-clients (default 1000). Each in-flight recursive query holds FDs. Reducing this cap limits recursive FD usage but causes SERVFAIL when the lower limit is reached.
  • Reduce the number of listen-on addresses. Each listen address consumes FDs for both UDP and TCP listeners. Consolidate where possible.
  • Disable query logging if enabled. Each open log file channel holds FDs, and query logging at high QPS generates excessive I/O that compounds the problem.

These are tradeoffs, not fixes. The right answer is to raise the FD limit to match the workload.

Address TCP-based attacks or transfer storms

If FD consumption is driven by abnormal TCP patterns:

  • A TCP SYN flood or slow-loris-style attack fills tcp-clients slots and exhausts FDs. Consider rate-limit configuration or upstream firewall rules to throttle TCP connection rates.
  • A zone transfer storm (many secondaries transferring simultaneously) is normal during refresh windows but can exhaust FDs on a busy server. Stagger secondary refresh schedules if possible.

Prevention

  • Set LimitNOFILE to at least 65536 in the systemd unit for every production named instance. Verify after deployment with grep "Max open files" /proc/$(pgrep -x named)/limits.
  • Monitor FD usage as a percentage of limit. Alert at 70%, page at 90%. Track the daily peak. If the daily peak exceeds 50% of the limit, plan to increase the limit or add capacity before the next traffic spike.
  • Remove the files option from named.conf on all BIND 9.18+ installations.
  • Test both UDP and TCP in health checks. UDP-only checks miss FD exhaustion because UDP listener sockets are persistent. TCP queries fail first when FDs are scarce.
  • Correlate FD usage with RecursClients and TCP connection count. These are the two largest FD consumers. If either trends upward, FD pressure follows.
  • Verify the limit after every deployment or config change. A package update or systemd unit change can silently reset LimitNOFILE to the default.

How Netdata helps

  • Per-second FD usage collection: Netdata’s apps or users plugin tracks open file descriptors per process, giving you a high-resolution view of FD consumption that catches spikes that 60-second polling misses.
  • FD limit correlation: Netdata collects process limits alongside usage, so you see the ratio directly rather than computing it manually.
  • BIND statistics channel integration: Netdata’s BIND collector pulls RecursClients, QrySERVFAIL, QuerySockFail, and socket statistics per second. When FD usage spikes, you can immediately see whether recursive clients, TCP connections, or socket failures drove it.
  • TCP connection state tracking: Netdata monitors ESTABLISHED, TIME_WAIT, and other TCP states on port 53, making it easy to distinguish transfer bursts from attacks.
  • Anomaly detection: ML-based anomaly flags on FD count and QuerySockFail catch slow FD leaks and gradual consumption growth before they hit the cliff edge.
  • Composite alerting: Correlate FD usage with SERVFAIL rate, recursive client count, and rndc response time to build alert conditions that distinguish FD exhaustion from other causes of query drops.