The only agent that thinks for itself

Autonomous Monitoring with self-learning AI built-in, operating independently across your entire stack.

Unlimited Metrics & Logs
Machine learning & MCP
5% CPU, 150MB RAM
3GB disk, >1 year retention
800+ integrations, zero config
Dashboards, alerts out of the box
> Discover Netdata Agents

Centralized metrics streaming and storage

Aggregate metrics from multiple agents into centralized Parent nodes for unified monitoring across your infrastructure.

Stream from unlimited agents
Long-term data retention
High availability clustering
Data replication & backup
Scalable architecture
Enterprise-grade security
> Learn about Parents

Fully managed cloud platform

Access your monitoring data from anywhere with our SaaS platform. No infrastructure to manage, automatic updates, and global availability.

Zero infrastructure management
99.9% uptime SLA
Global data centers
Automatic updates & patches
Enterprise SSO & RBAC
SOC2 & ISO certified
> Explore Netdata Cloud

Deploy Netdata Cloud in your infrastructure

Run the full Netdata Cloud platform on-premises for complete data sovereignty and compliance with your security policies.

Complete data sovereignty
Air-gapped deployment
Custom compliance controls
Private network integration
Dedicated support team
Kubernetes & Docker support
> Learn about Cloud On-Premises

Powerful, intuitive monitoring interface

Modern, responsive UI built for real-time troubleshooting with customizable dashboards and advanced visualization capabilities.

Real-time chart updates
Customizable dashboards
Dark & light themes
Advanced filtering & search
Responsive on all devices
Collaboration features
> Explore Netdata UI

Monitor on the go

Native iOS and Android apps bring full monitoring capabilities to your mobile device with real-time alerts and notifications.

iOS & Android apps
Push notifications
Touch-optimized interface
Offline data access
Biometric authentication
Widget support
> Download apps

The future of infrastructure observability

See our strategic direction across AI-native observability, full-stack signals, operational intelligence, and enterprise platform maturity.

AI-native observability
Full-stack signal coverage
Operational intelligence
Enterprise platform maturity
Agent releases every 6 weeks
Cloud continuous delivery
> Explore Product Roadmap

Best energy efficiency

True real-time per-second

100% automated zero config

Centralized observability

Multi-year retention

High availability built-in

Zero maintenance

Always up-to-date

Enterprise security

Complete data control

Air-gap ready

Compliance certified

Millisecond responsiveness

Infinite zoom & pan

Works on any device

Native performance

Instant alerts

Monitor anywhere

AI-native observability

Continuous delivery

Open source foundation

80% Faster Incident Resolution

AI-powered troubleshooting from detection, to root cause and blast radius identification, to reporting.

True Real-Time and Simple, even at Scale

Linearly and infinitely scalable full-stack observability, that can be deployed even mid-crisis.

90% Cost Reduction, Full Fidelity

Instead of centralizing the data, Netdata distributes the code, eliminating pipelines and complexity.

See and Map Your Entire Network

Live topology, flow analytics, and SNMP device and trap monitoring — unified with your full-stack observability.

Control Without Surrender

SOC 2 Type 2 certified with every metric kept on your infrastructure.

Integrations

800+ collectors and notification channels, auto-discovered and ready out of the box.

800+ data collectors
Auto-discovery & zero config
Cloud, infra, app protocols
Notifications out of the box
> Explore integrations
Real Results
46% Cost Reduction

Reduced monitoring costs by 46% while cutting staff overhead by 67%.

— Leonardo Antunez, Codyas

Zero Pipeline

No data shipping. No central storage costs. Query at the edge.

From Our Users
"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

No Query Language

Point-and-click troubleshooting. No PromQL, no LogQL, no learning curve.

Enterprise Ready
67% Less Staff, 46% Cost Cut

Enterprise efficiency without enterprise complexity—real ROI from day one.

— Leonardo Antunez, Codyas

SOC 2 Type 2 Certified

Zero data egress. Only metadata reaches the cloud. Your metrics stay on your infrastructure.

Full Coverage
800+ Collectors

Auto-discovered and configured. No manual setup required.

Any Notification Channel

Slack, PagerDuty, Teams, email, webhooks—all built-in.

Built for the People Who Get Paged

Because 3am alerts deserve instant answers, not hour-long hunts.

Every Industry Has Rules. We Master Them.

See how healthcare, finance, and government teams cut monitoring costs 90% while staying audit-ready.

Monitor Any Technology. Configure Nothing.

Install the agent. It already knows your stack.
From Our Users
"A Rare Unicorn"

Netdata gives more than you invest in it. A rare unicorn that obeys the Pareto rule.

— Eduard Porquet Mateu, TMB Barcelona

99% Downtime Reduction

Reduced website downtime by 99% and cloud bill by 30% using Netdata alerts.

— Falkland Islands Government

Real Savings
30% Cloud Cost Reduction

Optimized resource allocation based on Netdata alerts cut cloud spending by 30%.

— Falkland Islands Government

46% Cost Cut

Reduced monitoring staff by 67% while cutting operational costs by 46%.

— Codyas

Real Coverage
"Plugin for Everything"

Netdata has agent capacity or a plugin for everything, including Windows and Kubernetes.

— Eduard Porquet Mateu, TMB Barcelona

"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

Real Speed
Troubleshooting in 30 Seconds

From 2-3 minutes to 30 seconds—instant visibility into any node issue.

— Matthew Artist, Nodecraft

20% Downtime Reduction

20% less downtime and 40% budget optimization from out-of-the-box monitoring.

— Simon Beginn, LANCOM Systems

Pay per Node. Unlimited Everything Else.

One price per node. Unlimited metrics, logs, users, and retention. No per-GB surprises.

Free tier—forever
No metric limits or caps
Retention you control
Cancel anytime
> See pricing plans

What's Your Monitoring Really Costing You?

Most teams overpay by 40-60%. Let's find out why.

Expose hidden metric charges
Calculate tool consolidation
Customers report 30-67% savings
Results in under 60 seconds
> See what you're really paying

Your Infrastructure Is Unique. Let's Talk.

Because monitoring 10 nodes is different from monitoring 10,000.

On-prem & air-gapped deployment
Volume pricing & agreements
Architecture review for your scale
Compliance & security support
> Start a conversation

Monitoring That Sells Itself

Deploy in minutes. Impress clients in hours. Earn recurring revenue for years.

30-second live demos close deals
Zero config = zero support burden
Competitive margins & deal protection
Response in 48 hours
> Apply to partner

Per-Second Metrics at Homelab Prices

Same engine, same dashboards, same ML. Just priced for tinkerers.

Community: Free forever · 5 nodes · non-commercial
Homelab: $90/yr · unlimited nodes · fair usage
> Get the Homelab Plan

$1,000 Per Referral. Unlimited Referrals.

Your colleagues get 10% off. You get 10% commission. Everyone wins.

10% of subscriptions, up to $1,000 each
Track earnings inside Netdata Cloud
PayPal/Venmo payouts in 3-4 weeks
No caps, no complexity
> Get your referral link
Cost Proof
40% Budget Optimization

"Netdata's significant positive impact" — LANCOM Systems

Calculate Your Savings

Compare vs Datadog, Grafana, Dynatrace

Savings Proof
46% Cost Reduction

"Cut costs by 46%, staff by 67%" — Codyas

30% Cloud Bill Savings

"Reduced cloud bill by 30%" — Falkland Islands Gov

Enterprise Proof
"Better Than Combined Alternatives"

"Better observability with Netdata than combining other tools." — TMB Barcelona

Real Engineers, <24h Response

DPA, SLAs, on-prem, volume pricing

Why Partners Win
Demo Live Infrastructure

One command, 30 seconds, real data—no sandbox needed

Zero Tickets, High Margins

Auto-config + per-node pricing = predictable profit

Homelab Ready
Free Video Course

8-episode Netdata tutorial by LearnLinux.tv

76k+ GitHub Stars

3rd most starred monitoring project

Worth Recommending
Product That Delivers

Customers report 40-67% cost cuts, 99% downtime reduction

Zero Risk to Your Rep

Free tier lets them try before they buy

AI Support Assistant, Available 24/7

Nedi has access to all official documentation, source code, and resources. Ask any question about Netdata—responds in your language.

Deployment & configuration
Troubleshooting & sizing
Alerts & notifications
Evidence-based answers
> Ask Nedi now

Never Fight Fires Alone

Docs, community, and expert help—pick your path to resolution.

Learn.netdata.cloud docs
Discord, Forums, GitHub
Premium support available
> Get answers now

60 Seconds to First Dashboard

One command to install. Zero config. 850+ integrations documented.

Linux, Windows, K8s, Docker
Auto-discovers your stack
> Read our documentation

76,000+ Engineers Strong

615+ contributors. 1.5M daily downloads. One mission: simplify observability.

Per-Second. 90% Cheaper. Data Stays Home.

Side-by-side comparisons: costs, real-time granularity, and data sovereignty for every major tool.

See why teams switch from Datadog, Prometheus, Grafana, and more.

> Browse all comparisons
Edge-Native Observability, Born Open Source
Per-second visibility, ML on every metric, and data that never leaves your infrastructure.
Founded in 2016
615+ contributors worldwide
Remote-first, engineering-driven
Open source first
> Read our story
Promises We Publish—and Prove
12 principles backed by open code, independent validation, and measurable outcomes.
Open source, peer-reviewed
Zero config, instant value
Data sovereignty by design
Aligned pricing, no surprises
> See all 12 principles
Edge-Native, AI-Ready, 100% Open
76k+ stars. Full ML, AI, and automation—GPLv3+, not premium add-ons.
76,000+ GitHub stars
GPLv3+ licensed forever
ML on every metric, included
Zero vendor lock-in
> Explore our open source
Build Real-Time Observability for the World
Remote-first team shipping per-second monitoring with ML on every metric.
Remote-first, fully distributed
Open source (76k+ stars)
Challenging technical problems
Your code on millions of systems
> See open roles
Meet the Team Behind Netdata
Conferences, meetups, and tradeshows where you can see Netdata in action and talk to the engineers who build it.
Live demos and deep dives
Book 1-on-1 meetings
Talks and panel sessions
Event recaps and photos
> See all events
Talk to a Netdata Human in <24 Hours
Sales, partnerships, press, or professional services—real engineers, fast answers.
Discuss your observability needs
Pricing and volume discounts
Partnership opportunities
Media and press inquiries
> Book a conversation
Your Data. Your Rules.
On-prem data, cloud control plane, transparent terms.
Trust & Scale
76,000+ GitHub Stars

One of the most popular open-source monitoring projects

SOC 2 Type 2 Certified

Enterprise-grade security and compliance

Data Sovereignty

Your metrics stay on your infrastructure

Validated
University of Amsterdam

"Most energy-efficient monitoring solution" — ICSOC 2023, peer-reviewed

ADASTEC (Autonomous Driving)

"Doesn't miss alerts—mission-critical trust for safety software"

Community Stats
615+ Contributors

Global community improving monitoring for everyone

1.5M+ Downloads/Day

Trusted by teams worldwide

GPLv3+ Licensed

Free forever, fully open source agent

Why Join?
Remote-First

Work from anywhere, async-friendly culture

Impact at Scale

Your work helps millions of systems

$ guides / coredns / coredns-reload-failed ▌

Operations Guides

CoreDNS reload failed: config drift when the new Corefile never took effect

You pushed a Corefile change, nothing broke, and everyone moved on. Weeks later you discover the cluster is still running the old configuration. The metric that would have told you is coredns_reload_failed_total, and it was nonzero the whole time.

A failed CoreDNS reload is not an outage. The reload plugin polls the Corefile every 30 seconds (with jitter) and triggers a graceful reload when the SHA512 checksum of the file changes. If the new configuration has a syntax error, an invalid plugin directive, or hits a port conflict, CoreDNS logs the error, increments coredns_reload_failed_total, and keeps serving DNS with the old config. Queries keep flowing. The only thing that changed is that reality and your intent have quietly diverged.

That drift is the real incident. Every subsequent “we already fixed that in the Corefile” assumption is wrong until you reconcile what is loaded with what you think is loaded.

What this means

CoreDNS never partially applies a Corefile. A reload either succeeds atomically or the old configuration keeps running in full. There is no error returned to clients, no SERVFAIL spike, no crash. The failure is visible in exactly three places:

  • The counter coredns_reload_failed_total increments (typically once per poll cycle while the bad config sits on disk).
  • The CoreDNS log carries the parse or startup error for the rejected config.
  • The gauge coredns_reload_version_info{hash, value} records the SHA512 hash of the most recent reload attempt. It is updated to the attempted hash even when the reload fails, so a changed hash does not by itself prove the new config is running.

There is a nastier edge case documented upstream: a reload that changes a listener port closes the old listener before opening the new one. If opening the new port fails (for example, the port is already in use), the reload aborts but the old listener is already gone. DNS on port 53 keeps working while the health, ready, or metrics endpoint is permanently broken until the process restarts.

flowchart TD
  A[Corefile changed on disk] --> B[reload plugin polls every 30s with jitter]
  B --> C{SHA512 changed?}
  C -- no --> B
  C -- yes --> D[attempt graceful reload]
  D -- success --> E[new config live, cache flushed, version_info hash updates]
  D -- failure --> F[log error, reload_failed_total +1, old config keeps serving]
  F --> G[config drift: operator believes new config is live]

Any nonzero value of coredns_reload_failed_total is actionable. It means the running configuration does not match what someone intended to deploy.

Common causes

CauseWhat it looks likeFirst thing to check
Syntax or semantic error in the new CorefileCounter increments roughly once per poll cycle; parse error in logs naming the offending lineCoreDNS logs for the exact error text
Invalid plugin configuration (bad option, duplicate reload in one server block)Same as above; error mentions the pluginThe diff of the Corefile change
Listener port conflict during reloadReload fails; health, ready, or metrics endpoint stops responding while DNS still answersCurl each HTTP endpoint; find what owns the port
Change made only to an imported file on CoreDNS v1.6.0 or earlierCounter stays at zero, but coredns_reload_version_info never changes; reload plugin does not watch imported files on these versionsCoreDNS version; whether the edit was in an imported file
False alarm: poll interval not elapsed yetCounter is zero; hash updates 30 to 90 seconds after the file landsWait, then re-check the version hash
Version-specific reload wedge (v1.8.0 imported-file bug, upstream issue 5203)Every subsequent reload fails with “use of closed network connection”; next change can panic with “sync: negative WaitGroup counter”CoreDNS version and coredns_panics_total
Lameduck plus reload interaction (upstream issue 5471)Editing the config sends all pods unreadyReadiness endpoints during the change

Note on timing: in Kubernetes, the ConfigMap update has to propagate into the pod before the poll loop even sees it, and the Kubernetes documentation tells operators to allow up to two minutes for Corefile changes to take effect. Distinguish “reload has not happened yet” from “reload failed” before you start fixing things.

Quick checks

All read-only.

# 1. Check the reload failure counter
curl -s http://localhost:9153/metrics | grep '^coredns_reload_failed_total'

# 2. Check the hash of the most recent reload attempt (a failed reload still updates it)
curl -s http://localhost:9153/metrics | grep '^coredns_reload_version_info'

# 3. Find the reload error in the logs
kubectl logs -n kube-system -l k8s-app=kube-dns --tail=200 | grep -iE 'reload|corefile'

# 4. Read the intended configuration
kubectl get cm -n kube-system coredns -o yaml

# 5. Verify the HTTP endpoints survived (listener edge case)
curl -s -o /dev/null -w '%{http_code}\n' http://localhost:8080/health
curl -s -o /dev/null -w '%{http_code}\n' http://localhost:8181/ready
curl -s -o /dev/null -w '%{http_code}\n' http://localhost:9153/metrics

In Kubernetes, run the curl checks with kubectl exec against each CoreDNS pod. Check every replica: drift is frequently per-pod, because each pod polls and reloads independently.

How to diagnose it

  1. Confirm the failure is active, not historic. coredns_reload_failed_total is a counter and never decreases. Sample it twice, 60 to 90 seconds apart. If it is still incrementing, the on-disk config is being rejected on every poll cycle. If it is static, someone pushed a bad config in the past and the current file may be fine but was never loaded either.

  2. Read the error. The logs tell you exactly what CoreDNS rejected: a parse error with a line number, an unknown plugin option, or a startup failure such as a port bind. This is almost always enough to identify cause 1, 2, or 3 from the table above.

  3. Establish drift direction. Compare the intended Corefile (the ConfigMap) against coredns_reload_version_info on each pod. Record the hash, make a trivial known-good edit, wait up to two minutes, and record it again. If the hash never moves on a pod, that pod has not attempted a reload since it started. A moved hash is not proof of success: on a failed reload the gauge records the attempted new hash while the old config keeps running, so confirm with the failure counter and the logs.

  4. Rule out the timing false alarm. If the counter is zero and the change is recent, give it the full ConfigMap propagation plus poll interval (up to about two minutes) before concluding anything. If the hash then updates, there was never a failure.

  5. Check for the listener edge case. If the change touched any port (a new server block on a nonstandard port, or a moved health, ready, or prometheus directive), test every HTTP endpoint on every pod. A pod that answers DNS but returns nothing on :8080, :8181, or :9153 is in the wedged state: the old listener was closed and the new one never opened. The process cannot recover from this on its own.

  6. Check the version-specific failure modes. On v1.6.0 or earlier, edits to files pulled in with import do not trigger a reload at all; the counter stays zero and this looks like “nothing happened.” On v1.8.0, a failed reload triggered by an imported file can wedge the reload path: subsequent attempts fail with “use of closed network connection” and a later change can panic the process (“sync: negative WaitGroup counter”). Correlate coredns_panics_total and pod restarts with your config pushes. If you run lameduck for graceful shutdown, watch readiness closely during config changes.

Metrics and signals to monitor

SignalWhy it mattersWarning sign
coredns_reload_failed_totalDirect count of rejected reload attemptsAny nonzero value; repeated increments each poll cycle
coredns_reload_version_info{hash, value}SHA512 of the most recent reload attempt, successful or failedHash unchanged long after a config push, or hash moved without the failure counter moving; hash diverging between replicas
Cache hit ratio (coredns_cache_hits_total / coredns_cache_requests_total)A successful reload flushes the cache, producing a visible hit-ratio dip and forward-QPS spikeNo transient after an expected reload means the reload never happened; a transient with no config push means an unexpected reload
coredns_panics_totalPanics correlated with config changes indicate the wedged-reload class of bugAny increment near a Corefile change
Health self-check (coredns_health_request_failures_total, coredns_health_request_duration_seconds)Listener edge case degrades the health pathFailures or rising self-check latency after a reload attempt
/health, /ready, /metrics reachabilityThe wedge leaves DNS serving while endpoints are deadNon-200 or timeout on any of the three ports

Fixes

Fix the Corefile and push again

For syntax and plugin errors, correct the file and push it. The next poll cycle computes a new checksum, attempts the reload again, and on success the error loop stops. Confirm recovery by watching the version hash change and the cache transient occur. The counter does not reset; recovery is “stops incrementing,” not “returns to zero.” If the error complained about a duplicated directive, remember reload may only appear once per server block.

Recover from the listener port edge case

If a reload died midway through a port change, first free the conflicting port or pick a different one in the Corefile. The running process may still be missing its health or metrics listeners even after you fix the config, because the old listener was closed during the failed attempt. The reliable recovery is a pod restart. Treat that restart as a cache-flushing event (see below) and do one pod at a time.

When the change was never picked up

On v1.6.0 or earlier with import, make a trivial edit to the Corefile itself so the checksum changes, or plan an upgrade to v1.7.0 or later where imported files are tracked. On v1.8.0 with the imported-file wedge, upgrade; the reporter of upstream issue 5203 resolved it by moving to v1.9.0. If the process is already panic-looping, the pod restart is happening anyway.

Last resort: restart to force the config

If reloads keep failing and you cannot afford to wait, restarting the pod forces the new Corefile to load at startup. This is disruptive in one specific way: the cache comes back empty. Restarting all replicas at once converts your config fix into a cold-cache thundering herd against your upstreams. Delete pods one at a time, wait for readiness and cache warmup between them, and keep the change away from peak traffic if you can. See the cache collapse guide for why this matters.

Prevention

  • Alert on any reload failure. Any increase of coredns_reload_failed_total means running config and intended config differ. That is a ticket-worthy condition even though nothing is down.
  • Close the loop on every config push. After applying a Corefile change, automation should wait two minutes, then assert that coredns_reload_version_info changed on every replica and that no pod’s failure counter moved. A config push without this check is how drift survives for weeks.
  • Validate Corefiles before they ship. Render the final Corefile in CI and load it with a throwaway CoreDNS instance or staging pod before merging. Every rejected reload in production was a config that never got parsed anywhere first.
  • Do listener changes with a rollout, not a reload. Any edit that adds, removes, or moves a port should go through a controlled pod restart, because the reload path closes listeners before it knows the new ones will bind.
  • Track the version-specific bugs. Know whether you run anything at or below v1.6.0 (import blind spot), v1.8.0 (imported-file wedge), or with lameduck enabled (readiness interaction). Upgrade paths fix the first two.
  • Keep readiness on /ready. If a reload wedges a pod or a restart goes badly, correct readiness probes pull it out of the Service instead of letting it serve degraded answers. See the probe guide below.

How Netdata helps

  • Netdata scrapes each CoreDNS pod’s :9153 endpoint independently, so per-replica drift (one pod loaded the new config, the other rejected it) is visible instead of averaged away.
  • Alerting on any rate of change of coredns_reload_failed_total catches the failure on the first poll cycle, not when someone notices stale behavior weeks later.
  • Plotting coredns_reload_version_info alongside config push events closes the drift loop: you can see the hash move, or see it not move, per pod.
  • Correlating reload events with the cache hit ratio distinguishes “reload happened” (hit-ratio dip, forward-QPS spike) from “reload silently never happened” (flat lines after a push).
  • coredns_panics_total and pod restarts on the same timeline as config changes surface the wedged-reload bug class before it takes pods down.
  • The health self-check metrics (coredns_health_request_failures_total, coredns_health_request_duration_seconds) reveal the listener edge case where DNS still answers but the health path is broken.