The only agent that thinks for itself

Autonomous Monitoring with self-learning AI built-in, operating independently across your entire stack.

Unlimited Metrics & Logs
Machine learning & MCP
5% CPU, 150MB RAM
3GB disk, >1 year retention
800+ integrations, zero config
Dashboards, alerts out of the box
> Discover Netdata Agents

Centralized metrics streaming and storage

Aggregate metrics from multiple agents into centralized Parent nodes for unified monitoring across your infrastructure.

Stream from unlimited agents
Long-term data retention
High availability clustering
Data replication & backup
Scalable architecture
Enterprise-grade security
> Learn about Parents

Fully managed cloud platform

Access your monitoring data from anywhere with our SaaS platform. No infrastructure to manage, automatic updates, and global availability.

Zero infrastructure management
99.9% uptime SLA
Global data centers
Automatic updates & patches
Enterprise SSO & RBAC
SOC2 & ISO certified
> Explore Netdata Cloud

Deploy Netdata Cloud in your infrastructure

Run the full Netdata Cloud platform on-premises for complete data sovereignty and compliance with your security policies.

Complete data sovereignty
Air-gapped deployment
Custom compliance controls
Private network integration
Dedicated support team
Kubernetes & Docker support
> Learn about Cloud On-Premises

Powerful, intuitive monitoring interface

Modern, responsive UI built for real-time troubleshooting with customizable dashboards and advanced visualization capabilities.

Real-time chart updates
Customizable dashboards
Dark & light themes
Advanced filtering & search
Responsive on all devices
Collaboration features
> Explore Netdata UI

Monitor on the go

Native iOS and Android apps bring full monitoring capabilities to your mobile device with real-time alerts and notifications.

iOS & Android apps
Push notifications
Touch-optimized interface
Offline data access
Biometric authentication
Widget support
> Download apps

The future of infrastructure observability

See our strategic direction across AI-native observability, full-stack signals, operational intelligence, and enterprise platform maturity.

AI-native observability
Full-stack signal coverage
Operational intelligence
Enterprise platform maturity
Agent releases every 6 weeks
Cloud continuous delivery
> Explore Product Roadmap

Best energy efficiency

True real-time per-second

100% automated zero config

Centralized observability

Multi-year retention

High availability built-in

Zero maintenance

Always up-to-date

Enterprise security

Complete data control

Air-gap ready

Compliance certified

Millisecond responsiveness

Infinite zoom & pan

Works on any device

Native performance

Instant alerts

Monitor anywhere

AI-native observability

Continuous delivery

Open source foundation

80% Faster Incident Resolution

AI-powered troubleshooting from detection, to root cause and blast radius identification, to reporting.

True Real-Time and Simple, even at Scale

Linearly and infinitely scalable full-stack observability, that can be deployed even mid-crisis.

90% Cost Reduction, Full Fidelity

Instead of centralizing the data, Netdata distributes the code, eliminating pipelines and complexity.

See and Map Your Entire Network

Live topology, flow analytics, and SNMP device and trap monitoring — unified with your full-stack observability.

Control Without Surrender

SOC 2 Type 2 certified with every metric kept on your infrastructure.

Integrations

800+ collectors and notification channels, auto-discovered and ready out of the box.

800+ data collectors
Auto-discovery & zero config
Cloud, infra, app protocols
Notifications out of the box
> Explore integrations
Real Results
46% Cost Reduction

Reduced monitoring costs by 46% while cutting staff overhead by 67%.

— Leonardo Antunez, Codyas

Zero Pipeline

No data shipping. No central storage costs. Query at the edge.

From Our Users
"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

No Query Language

Point-and-click troubleshooting. No PromQL, no LogQL, no learning curve.

Enterprise Ready
67% Less Staff, 46% Cost Cut

Enterprise efficiency without enterprise complexity—real ROI from day one.

— Leonardo Antunez, Codyas

SOC 2 Type 2 Certified

Zero data egress. Only metadata reaches the cloud. Your metrics stay on your infrastructure.

Full Coverage
800+ Collectors

Auto-discovered and configured. No manual setup required.

Any Notification Channel

Slack, PagerDuty, Teams, email, webhooks—all built-in.

Built for the People Who Get Paged

Because 3am alerts deserve instant answers, not hour-long hunts.

Every Industry Has Rules. We Master Them.

See how healthcare, finance, and government teams cut monitoring costs 90% while staying audit-ready.

Monitor Any Technology. Configure Nothing.

Install the agent. It already knows your stack.
From Our Users
"A Rare Unicorn"

Netdata gives more than you invest in it. A rare unicorn that obeys the Pareto rule.

— Eduard Porquet Mateu, TMB Barcelona

99% Downtime Reduction

Reduced website downtime by 99% and cloud bill by 30% using Netdata alerts.

— Falkland Islands Government

Real Savings
30% Cloud Cost Reduction

Optimized resource allocation based on Netdata alerts cut cloud spending by 30%.

— Falkland Islands Government

46% Cost Cut

Reduced monitoring staff by 67% while cutting operational costs by 46%.

— Codyas

Real Coverage
"Plugin for Everything"

Netdata has agent capacity or a plugin for everything, including Windows and Kubernetes.

— Eduard Porquet Mateu, TMB Barcelona

"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

Real Speed
Troubleshooting in 30 Seconds

From 2-3 minutes to 30 seconds—instant visibility into any node issue.

— Matthew Artist, Nodecraft

20% Downtime Reduction

20% less downtime and 40% budget optimization from out-of-the-box monitoring.

— Simon Beginn, LANCOM Systems

Pay per Node. Unlimited Everything Else.

One price per node. Unlimited metrics, logs, users, and retention. No per-GB surprises.

Free tier—forever
No metric limits or caps
Retention you control
Cancel anytime
> See pricing plans

What's Your Monitoring Really Costing You?

Most teams overpay by 40-60%. Let's find out why.

Expose hidden metric charges
Calculate tool consolidation
Customers report 30-67% savings
Results in under 60 seconds
> See what you're really paying

Your Infrastructure Is Unique. Let's Talk.

Because monitoring 10 nodes is different from monitoring 10,000.

On-prem & air-gapped deployment
Volume pricing & agreements
Architecture review for your scale
Compliance & security support
> Start a conversation

Monitoring That Sells Itself

Deploy in minutes. Impress clients in hours. Earn recurring revenue for years.

30-second live demos close deals
Zero config = zero support burden
Competitive margins & deal protection
Response in 48 hours
> Apply to partner

Per-Second Metrics at Homelab Prices

Same engine, same dashboards, same ML. Just priced for tinkerers.

Community: Free forever · 5 nodes · non-commercial
Homelab: $90/yr · unlimited nodes · fair usage
> Get the Homelab Plan

$1,000 Per Referral. Unlimited Referrals.

Your colleagues get 10% off. You get 10% commission. Everyone wins.

10% of subscriptions, up to $1,000 each
Track earnings inside Netdata Cloud
PayPal/Venmo payouts in 3-4 weeks
No caps, no complexity
> Get your referral link
Cost Proof
40% Budget Optimization

"Netdata's significant positive impact" — LANCOM Systems

Calculate Your Savings

Compare vs Datadog, Grafana, Dynatrace

Savings Proof
46% Cost Reduction

"Cut costs by 46%, staff by 67%" — Codyas

30% Cloud Bill Savings

"Reduced cloud bill by 30%" — Falkland Islands Gov

Enterprise Proof
"Better Than Combined Alternatives"

"Better observability with Netdata than combining other tools." — TMB Barcelona

Real Engineers, <24h Response

DPA, SLAs, on-prem, volume pricing

Why Partners Win
Demo Live Infrastructure

One command, 30 seconds, real data—no sandbox needed

Zero Tickets, High Margins

Auto-config + per-node pricing = predictable profit

Homelab Ready
Free Video Course

8-episode Netdata tutorial by LearnLinux.tv

76k+ GitHub Stars

3rd most starred monitoring project

Worth Recommending
Product That Delivers

Customers report 40-67% cost cuts, 99% downtime reduction

Zero Risk to Your Rep

Free tier lets them try before they buy

AI Support Assistant, Available 24/7

Nedi has access to all official documentation, source code, and resources. Ask any question about Netdata—responds in your language.

Deployment & configuration
Troubleshooting & sizing
Alerts & notifications
Evidence-based answers
> Ask Nedi now

Never Fight Fires Alone

Docs, community, and expert help—pick your path to resolution.

Learn.netdata.cloud docs
Discord, Forums, GitHub
Premium support available
> Get answers now

60 Seconds to First Dashboard

One command to install. Zero config. 850+ integrations documented.

Linux, Windows, K8s, Docker
Auto-discovers your stack
> Read our documentation

76,000+ Engineers Strong

615+ contributors. 1.5M daily downloads. One mission: simplify observability.

Per-Second. 90% Cheaper. Data Stays Home.

Side-by-side comparisons: costs, real-time granularity, and data sovereignty for every major tool.

See why teams switch from Datadog, Prometheus, Grafana, and more.

> Browse all comparisons
Edge-Native Observability, Born Open Source
Per-second visibility, ML on every metric, and data that never leaves your infrastructure.
Founded in 2016
615+ contributors worldwide
Remote-first, engineering-driven
Open source first
> Read our story
Promises We Publish—and Prove
12 principles backed by open code, independent validation, and measurable outcomes.
Open source, peer-reviewed
Zero config, instant value
Data sovereignty by design
Aligned pricing, no surprises
> See all 12 principles
Edge-Native, AI-Ready, 100% Open
76k+ stars. Full ML, AI, and automation—GPLv3+, not premium add-ons.
76,000+ GitHub stars
GPLv3+ licensed forever
ML on every metric, included
Zero vendor lock-in
> Explore our open source
Build Real-Time Observability for the World
Remote-first team shipping per-second monitoring with ML on every metric.
Remote-first, fully distributed
Open source (76k+ stars)
Challenging technical problems
Your code on millions of systems
> See open roles
Meet the Team Behind Netdata
Conferences, meetups, and tradeshows where you can see Netdata in action and talk to the engineers who build it.
Live demos and deep dives
Book 1-on-1 meetings
Talks and panel sessions
Event recaps and photos
> See all events
Talk to a Netdata Human in <24 Hours
Sales, partnerships, press, or professional services—real engineers, fast answers.
Discuss your observability needs
Pricing and volume discounts
Partnership opportunities
Media and press inquiries
> Book a conversation
Your Data. Your Rules.
On-prem data, cloud control plane, transparent terms.
Trust & Scale
76,000+ GitHub Stars

One of the most popular open-source monitoring projects

SOC 2 Type 2 Certified

Enterprise-grade security and compliance

Data Sovereignty

Your metrics stay on your infrastructure

Validated
University of Amsterdam

"Most energy-efficient monitoring solution" — ICSOC 2023, peer-reviewed

ADASTEC (Autonomous Driving)

"Doesn't miss alerts—mission-critical trust for safety software"

Community Stats
615+ Contributors

Global community improving monitoring for everyone

1.5M+ Downloads/Day

Trusted by teams worldwide

GPLv3+ Licensed

Free forever, fully open source agent

Why Join?
Remote-First

Work from anywhere, async-friendly culture

Impact at Scale

Your work helps millions of systems

$ guides / bind-dns / bind-dns-cold-cache-warming ▌

Operations Guides

BIND cold cache after restart: the warming storm and elevated upstream load

After a named restart, the recursive cache is empty. Every incoming query is a cache miss. On a resolver doing 50,000 qps, that means at least 50,000 outbound recursive queries per second to upstream authoritative servers – a 10x increase over steady-state outbound volume. This is the cache-warming storm, and it lasts 30 to 60 minutes.

Three things happen simultaneously: upstream load spikes, the recursive-clients table fills toward its limit, and cache hit ratio starts at zero and climbs. For the first 30 to 60 seconds, some queries may return SERVFAIL while root priming completes and authoritative zones finish loading. All of this is expected behavior.

The operational problem is alerting. Monitoring that does not account for cold starts will page on every restart: cache hit ratio alerts fire at 0%, SERVFAIL alerts fire during root priming, and recursive client alerts fire as the table fills. The fix is uptime gating on every alert that depends on cache state.

What happens during a cold start

BIND has no mechanism to persist its cache to disk and reload it on startup. rndc dumpdb produces a human-readable dump for manual inspection, but there is no corresponding load operation. Every restart produces a full cold cache.

The cold-start sequence proceeds in phases:

  1. Process startup and zone loading (seconds to minutes). named starts, reads configuration, and loads authoritative zone data into memory. For servers with many large zones, validation and loading takes minutes. During this window, the server may not respond to queries at all.

  2. Root priming (first 30 to 60 seconds). Root hint priming happens on demand when the first recursive query arrives that requires contacting a root nameserver. BIND forwards the client query to a root hint server while priming proceeds in parallel. Until priming completes, some queries return SERVFAIL.

  3. RTT probing (seconds to minutes). After startup, BIND probes upstream nameservers to estimate round-trip times for server selection. Server selection is suboptimal during this phase and resolution latency is temporarily higher than steady state.

  4. Cache warming (30 to 60 minutes). The cache fills as queries arrive and get resolved. Hit ratio climbs from zero toward steady-state levels. Outbound recursive query volume is far higher than normal, placing elevated load on upstream authoritative servers.

flowchart TD
    A["named restart"] --> B["Zone load\n(seconds to minutes)"]
    B --> C["Root priming\n(first 30-60s)"]
    C --> D["RTT probing\n(seconds to minutes)"]
    D --> E["Cache warming\n(30-60 min)"]
    C -.->|"SERVFAIL possible"| F["Brief query failures\nuntil priming completes"]
    E -.->|"Elevated upstream load"| G["Outbound QPS\napproaches inbound"]
    E -.->|"Slot pressure"| H["RecursClients fills\ntoward limit"]
    E --> I["Steady state:\nhit ratio normalizes"]

The warming storm: why restarts hit upstream hard

In steady state, a healthy resolver answers 80 to 95% of queries from cache, generating only 5 to 20% of inbound volume as outbound recursive queries. After a restart, that ratio inverts: outbound volume approaches or equals inbound volume because every query is a miss.

On a resolver doing 50,000 qps with a steady-state hit ratio of 90%, normal outbound volume is approximately 5,000 qps upstream. After restart, outbound volume jumps toward 50,000 qps. That is a 10x amplification of upstream load, and it persists until the cache reaches working-set coverage.

The amplification is worse for names with complex resolution chains. Domains with long CNAME chains and out-of-bailiwick nameserver names require multiple sequential queries per resolution. A single user query for a domain with a four-step CNAME chain spread across multiple TLDs can require dozens of upstream queries when the cache is cold. Each upstream query occupies a recursive-client slot for the duration of its resolution.

This is why RecursClients spikes after restart. Every cache miss generates one or more in-flight recursive queries. With recursive-clients defaulting to 1000, a high-traffic resolver can approach the soft quota (900) within seconds of startup. At the hard limit (1000), all new recursive queries receive SERVFAIL.

What is normal after restart

Several behaviors that look alarming are expected during the cold-start window.

SERVFAIL in the first 30 to 60 seconds. Root priming, zone loading, and managed-keys validation all happen at startup. Some queries fail during this window. It resolves on its own as priming completes.

Cache hit ratio at zero. The cache is empty. Hit ratio starts at zero and climbs over 30 to 60 minutes as the working set populates.

Elevated recursive client count. Every cache miss generates recursive traffic. On a high-traffic resolver, RecursClients approaches or temporarily exceeds the soft quota. This is the warming storm in progress, not a sign of upstream failure.

Suboptimal upstream server selection. BIND’s RTT estimates reset on restart. During the probing phase, queries may route to slower upstream nameservers until BIND learns which are fastest.

Memory growth. RSS climbs for 30 to 60 minutes as the cache fills. RSS should stabilize after the working set is cached. BIND’s internal allocator does not return freed memory to the OS efficiently, so RSS may remain elevated even after cache contents cycle.

What is NOT normal after restart

SERVFAIL persisting beyond 60 seconds. If SERVFAIL rate remains elevated after the first minute, something beyond root priming is wrong. Check DNSSEC managed-keys (rndc managed-keys status), zone load status in logs, and upstream reachability.

Cache hit ratio not recovering. If hit ratio stays near zero after 30 minutes, the cache is not warming. This indicates either a workload dominated by random subdomains (water torture attack) or a systematic problem with upstream resolution.

RecursClients pinned at the limit. Brief elevation during cache warming is expected. Sustained saturation beyond 5 minutes, combined with rising SERVFAIL, indicates a recursive resolution cascade that is not self-correcting.

Zones not loading. named starts successfully even if individual zones fail to load. Check logs for zone load errors after every restart. A server can report “running” via rndc status while serving REFUSED or SERVFAIL for specific zones.

Why alerts need uptime gates

Every alert that depends on cache state must gate on uptime. Without an uptime gate, every restart triggers a cascade of false pages.

Alert conditionUptime gateRationale
SERVFAIL rate above 1%Greater than 300 secondsRoot priming causes SERVFAIL in first 30-60s
RecursClients above 90% of limitGreater than 300 secondsCold cache fills recursive slots immediately after restart
DNSSEC ValFail spikeGreater than 300 secondsKey initialization produces transient validation noise
Cache hit ratio below baselineGreater than 1800 secondsCache warming takes 30-60 minutes to reach steady state

The 300-second gate covers root priming and RTT probing. The 1800-second gate covers cache warming. For planned restarts, consider suppressing alerts during a maintenance window rather than relying solely on uptime gates. The uptime gate is the safety net for unplanned restarts, OOM kills, and crash recovery.

Handling restarts on high-traffic resolvers

On resolvers handling significant traffic, the warming storm can cause real upstream impact. Several strategies reduce the blast radius.

Stagger restarts in multi-resolver or anycast setups. Restarting all nodes simultaneously multiplies the upstream load. Restart nodes sequentially, allowing each to warm before restarting the next.

Consider pre-warming. There is no built-in cache pre-warming mechanism, but a script that sends a representative sample of production queries to the freshly restarted resolver can accelerate cache population. This does not eliminate the warming storm but shortens it by seeding the cache with the most common names. Rate-limit the pre-warm queries to avoid amplifying the upstream load you are trying to mitigate.

Monitor RecursClients during warming. If recursive clients hit the hard limit during warming, the resolver returns SERVFAIL for all new recursive queries. Temporarily increasing recursive-clients (with sufficient FD and memory headroom) can provide breathing room during the warm-up phase.

Reset statistics after warming. On BIND 9.21 and later, rndc reset-stats clears high-water counters, allowing clean measurement of steady-state performance without the cold-start peak distorting the baseline.

Signals to watch during cache warming

SignalWhy it mattersExpected pattern during warming
Cache hit ratioCore warming indicatorStarts at 0%, climbs over 30-60 min
RecursClientsPressure on recursive slotsSpikes immediately, declines as cache fills
Outbound query rateUpstream load amplifierApproaches inbound rate, then drops as hit ratio climbs
SERVFAIL rateDistinguishes normal priming from real failureBrief spike in first 60s, then near-zero
Process RSSMemory growth from cache populationClimbs for 30-60 min, then stabilizes
QryRTT distributionUpstream server selection qualitySuboptimal initially, improves as RTT estimates populate

The key correlation: if cache hit ratio is climbing and RecursClients is declining, the warming storm is progressing normally. If either stalls, investigate. A restart event followed by rising cache misses, climbing RecursClients, and a brief SERVFAIL spike is normal cold-start behavior. The same signals without recovery after 30 minutes point to a different problem.

How Netdata helps

Netdata’s per-second metrics make the warming curve visible in real time. During a cold start, correlate:

  • Cache hit ratio and cache misses per second: per-second resolution shows the warming curve in detail, making it clear when the cache has reached steady state versus stalled at a low level.
  • Recursive clients as a percentage of limit: tracking RecursClients against the configured recursive-clients limit shows whether the warming storm is approaching the circuit breaker threshold.
  • SERVFAIL rate correlated with process uptime: combining SERVFAIL metrics with uptime distinguishes root-priming noise in the first 60 seconds from real resolution failures that persist.
  • Outbound query rate: the ratio of outbound to inbound queries directly shows cache warming progress and upstream load amplification.
  • QryRTT distribution per view: RTT buckets show when BIND’s server selection has converged on optimal upstream nameservers after the post-restart probing phase.

These signals are most useful on a single timeline. The warming storm has a recognizable shape: cache misses spike, RecursClients surges, SERVFAIL blips briefly, then all three recover as the cache fills. Deviation from that shape is the diagnostic signal.