The only agent that thinks for itself

Autonomous Monitoring with self-learning AI built-in, operating independently across your entire stack.

Unlimited Metrics & Logs
Machine learning & MCP
5% CPU, 150MB RAM
3GB disk, >1 year retention
800+ integrations, zero config
Dashboards, alerts out of the box
> Discover Netdata Agents

Centralized metrics streaming and storage

Aggregate metrics from multiple agents into centralized Parent nodes for unified monitoring across your infrastructure.

Stream from unlimited agents
Long-term data retention
High availability clustering
Data replication & backup
Scalable architecture
Enterprise-grade security
> Learn about Parents

Fully managed cloud platform

Access your monitoring data from anywhere with our SaaS platform. No infrastructure to manage, automatic updates, and global availability.

Zero infrastructure management
99.9% uptime SLA
Global data centers
Automatic updates & patches
Enterprise SSO & RBAC
SOC2 & ISO certified
> Explore Netdata Cloud

Deploy Netdata Cloud in your infrastructure

Run the full Netdata Cloud platform on-premises for complete data sovereignty and compliance with your security policies.

Complete data sovereignty
Air-gapped deployment
Custom compliance controls
Private network integration
Dedicated support team
Kubernetes & Docker support
> Learn about Cloud On-Premises

Powerful, intuitive monitoring interface

Modern, responsive UI built for real-time troubleshooting with customizable dashboards and advanced visualization capabilities.

Real-time chart updates
Customizable dashboards
Dark & light themes
Advanced filtering & search
Responsive on all devices
Collaboration features
> Explore Netdata UI

Monitor on the go

Native iOS and Android apps bring full monitoring capabilities to your mobile device with real-time alerts and notifications.

iOS & Android apps
Push notifications
Touch-optimized interface
Offline data access
Biometric authentication
Widget support
> Download apps

The future of infrastructure observability

See our strategic direction across AI-native observability, full-stack signals, operational intelligence, and enterprise platform maturity.

AI-native observability
Full-stack signal coverage
Operational intelligence
Enterprise platform maturity
Agent releases every 6 weeks
Cloud continuous delivery
> Explore Product Roadmap

Best energy efficiency

True real-time per-second

100% automated zero config

Centralized observability

Multi-year retention

High availability built-in

Zero maintenance

Always up-to-date

Enterprise security

Complete data control

Air-gap ready

Compliance certified

Millisecond responsiveness

Infinite zoom & pan

Works on any device

Native performance

Instant alerts

Monitor anywhere

AI-native observability

Continuous delivery

Open source foundation

80% Faster Incident Resolution

AI-powered troubleshooting from detection, to root cause and blast radius identification, to reporting.

True Real-Time and Simple, even at Scale

Linearly and infinitely scalable full-stack observability, that can be deployed even mid-crisis.

90% Cost Reduction, Full Fidelity

Instead of centralizing the data, Netdata distributes the code, eliminating pipelines and complexity.

See and Map Your Entire Network

Live topology, flow analytics, and SNMP device and trap monitoring — unified with your full-stack observability.

Control Without Surrender

SOC 2 Type 2 certified with every metric kept on your infrastructure.

Integrations

800+ collectors and notification channels, auto-discovered and ready out of the box.

800+ data collectors
Auto-discovery & zero config
Cloud, infra, app protocols
Notifications out of the box
> Explore integrations
Real Results
46% Cost Reduction

Reduced monitoring costs by 46% while cutting staff overhead by 67%.

— Leonardo Antunez, Codyas

Zero Pipeline

No data shipping. No central storage costs. Query at the edge.

From Our Users
"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

No Query Language

Point-and-click troubleshooting. No PromQL, no LogQL, no learning curve.

Enterprise Ready
67% Less Staff, 46% Cost Cut

Enterprise efficiency without enterprise complexity—real ROI from day one.

— Leonardo Antunez, Codyas

SOC 2 Type 2 Certified

Zero data egress. Only metadata reaches the cloud. Your metrics stay on your infrastructure.

Full Coverage
800+ Collectors

Auto-discovered and configured. No manual setup required.

Any Notification Channel

Slack, PagerDuty, Teams, email, webhooks—all built-in.

Built for the People Who Get Paged

Because 3am alerts deserve instant answers, not hour-long hunts.

Every Industry Has Rules. We Master Them.

See how healthcare, finance, and government teams cut monitoring costs 90% while staying audit-ready.

Monitor Any Technology. Configure Nothing.

Install the agent. It already knows your stack.
From Our Users
"A Rare Unicorn"

Netdata gives more than you invest in it. A rare unicorn that obeys the Pareto rule.

— Eduard Porquet Mateu, TMB Barcelona

99% Downtime Reduction

Reduced website downtime by 99% and cloud bill by 30% using Netdata alerts.

— Falkland Islands Government

Real Savings
30% Cloud Cost Reduction

Optimized resource allocation based on Netdata alerts cut cloud spending by 30%.

— Falkland Islands Government

46% Cost Cut

Reduced monitoring staff by 67% while cutting operational costs by 46%.

— Codyas

Real Coverage
"Plugin for Everything"

Netdata has agent capacity or a plugin for everything, including Windows and Kubernetes.

— Eduard Porquet Mateu, TMB Barcelona

"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

Real Speed
Troubleshooting in 30 Seconds

From 2-3 minutes to 30 seconds—instant visibility into any node issue.

— Matthew Artist, Nodecraft

20% Downtime Reduction

20% less downtime and 40% budget optimization from out-of-the-box monitoring.

— Simon Beginn, LANCOM Systems

Pay per Node. Unlimited Everything Else.

One price per node. Unlimited metrics, logs, users, and retention. No per-GB surprises.

Free tier—forever
No metric limits or caps
Retention you control
Cancel anytime
> See pricing plans

What's Your Monitoring Really Costing You?

Most teams overpay by 40-60%. Let's find out why.

Expose hidden metric charges
Calculate tool consolidation
Customers report 30-67% savings
Results in under 60 seconds
> See what you're really paying

Your Infrastructure Is Unique. Let's Talk.

Because monitoring 10 nodes is different from monitoring 10,000.

On-prem & air-gapped deployment
Volume pricing & agreements
Architecture review for your scale
Compliance & security support
> Start a conversation

Monitoring That Sells Itself

Deploy in minutes. Impress clients in hours. Earn recurring revenue for years.

30-second live demos close deals
Zero config = zero support burden
Competitive margins & deal protection
Response in 48 hours
> Apply to partner

Per-Second Metrics at Homelab Prices

Same engine, same dashboards, same ML. Just priced for tinkerers.

Community: Free forever · 5 nodes · non-commercial
Homelab: $90/yr · unlimited nodes · fair usage
> Get the Homelab Plan

$1,000 Per Referral. Unlimited Referrals.

Your colleagues get 10% off. You get 10% commission. Everyone wins.

10% of subscriptions, up to $1,000 each
Track earnings inside Netdata Cloud
PayPal/Venmo payouts in 3-4 weeks
No caps, no complexity
> Get your referral link
Cost Proof
40% Budget Optimization

"Netdata's significant positive impact" — LANCOM Systems

Calculate Your Savings

Compare vs Datadog, Grafana, Dynatrace

Savings Proof
46% Cost Reduction

"Cut costs by 46%, staff by 67%" — Codyas

30% Cloud Bill Savings

"Reduced cloud bill by 30%" — Falkland Islands Gov

Enterprise Proof
"Better Than Combined Alternatives"

"Better observability with Netdata than combining other tools." — TMB Barcelona

Real Engineers, <24h Response

DPA, SLAs, on-prem, volume pricing

Why Partners Win
Demo Live Infrastructure

One command, 30 seconds, real data—no sandbox needed

Zero Tickets, High Margins

Auto-config + per-node pricing = predictable profit

Homelab Ready
Free Video Course

8-episode Netdata tutorial by LearnLinux.tv

76k+ GitHub Stars

3rd most starred monitoring project

Worth Recommending
Product That Delivers

Customers report 40-67% cost cuts, 99% downtime reduction

Zero Risk to Your Rep

Free tier lets them try before they buy

AI Support Assistant, Available 24/7

Nedi has access to all official documentation, source code, and resources. Ask any question about Netdata—responds in your language.

Deployment & configuration
Troubleshooting & sizing
Alerts & notifications
Evidence-based answers
> Ask Nedi now

Never Fight Fires Alone

Docs, community, and expert help—pick your path to resolution.

Learn.netdata.cloud docs
Discord, Forums, GitHub
Premium support available
> Get answers now

60 Seconds to First Dashboard

One command to install. Zero config. 850+ integrations documented.

Linux, Windows, K8s, Docker
Auto-discovers your stack
> Read our documentation

76,000+ Engineers Strong

615+ contributors. 1.5M daily downloads. One mission: simplify observability.

Per-Second. 90% Cheaper. Data Stays Home.

Side-by-side comparisons: costs, real-time granularity, and data sovereignty for every major tool.

See why teams switch from Datadog, Prometheus, Grafana, and more.

> Browse all comparisons
Edge-Native Observability, Born Open Source
Per-second visibility, ML on every metric, and data that never leaves your infrastructure.
Founded in 2016
615+ contributors worldwide
Remote-first, engineering-driven
Open source first
> Read our story
Promises We Publish—and Prove
12 principles backed by open code, independent validation, and measurable outcomes.
Open source, peer-reviewed
Zero config, instant value
Data sovereignty by design
Aligned pricing, no surprises
> See all 12 principles
Edge-Native, AI-Ready, 100% Open
76k+ stars. Full ML, AI, and automation—GPLv3+, not premium add-ons.
76,000+ GitHub stars
GPLv3+ licensed forever
ML on every metric, included
Zero vendor lock-in
> Explore our open source
Build Real-Time Observability for the World
Remote-first team shipping per-second monitoring with ML on every metric.
Remote-first, fully distributed
Open source (76k+ stars)
Challenging technical problems
Your code on millions of systems
> See open roles
Meet the Team Behind Netdata
Conferences, meetups, and tradeshows where you can see Netdata in action and talk to the engineers who build it.
Live demos and deep dives
Book 1-on-1 meetings
Talks and panel sessions
Event recaps and photos
> See all events
Talk to a Netdata Human in <24 Hours
Sales, partnerships, press, or professional services—real engineers, fast answers.
Discuss your observability needs
Pricing and volume discounts
Partnership opportunities
Media and press inquiries
> Book a conversation
Your Data. Your Rules.
On-prem data, cloud control plane, transparent terms.
Trust & Scale
76,000+ GitHub Stars

One of the most popular open-source monitoring projects

SOC 2 Type 2 Certified

Enterprise-grade security and compliance

Data Sovereignty

Your metrics stay on your infrastructure

Validated
University of Amsterdam

"Most energy-efficient monitoring solution" — ICSOC 2023, peer-reviewed

ADASTEC (Autonomous Driving)

"Doesn't miss alerts—mission-critical trust for safety software"

Community Stats
615+ Contributors

Global community improving monitoring for everyone

1.5M+ Downloads/Day

Trusted by teams worldwide

GPLv3+ Licensed

Free forever, fully open source agent

Why Join?
Remote-First

Work from anywhere, async-friendly culture

Impact at Scale

Your work helps millions of systems

$ guides / coredns / coredns-nxdomain-vs-servfail ▌

Operations Guides

CoreDNS NXDOMAIN vs SERVFAIL: why alerting on the wrong one buries real incidents

Your CoreDNS alert fired again. The “DNS error rate” panel shows a wall of red, so you open it, see a pile of NXDOMAIN responses, and silence the alert. Meanwhile, a genuine upstream failure is producing SERVFAIL responses somewhere in that same wall of red, and nobody will notice until an application team opens a ticket.

This is the most common CoreDNS alerting mistake: treating all non-NOERROR responses as one bucket. NXDOMAIN and SERVFAIL look similar in a naive error counter, but they mean opposite things. NXDOMAIN is the server correctly reporting that a name does not exist. SERVFAIL is the server admitting it failed to answer at all. Alerting on the sum of both guarantees that the noise (NXDOMAIN, constant and expected in Kubernetes) drowns the signal (SERVFAIL, rare and always actionable).

What this means

DNS response codes (RCODEs) tell the client what happened to its query. Four of them matter here:

  • NOERROR: the query was answered successfully. This includes NODATA (the name exists but has no record of the requested type), which is also a successful answer.
  • NXDOMAIN: the name does not exist. The resolver looked, confirmed the name is not there, and said so. This is a correct, authoritative answer, not an operational error.
  • SERVFAIL: the server failed to process the query. Upstream unreachable, plugin failure, API connectivity loss, configuration error. This is an actual failure of the resolution path.
  • REFUSED: the server explicitly rejected the query (no matching server block, ACL rejection, or forward max_concurrent limit hit). Also a real failure, usually misconfiguration or capacity.

CoreDNS exposes all of these on one counter family, coredns_dns_responses_total, with an rcode label. The mistake is writing an alert like “rate of responses where rcode != NOERROR.” In a Kubernetes cluster, that numerator is dominated by NXDOMAIN, so the alert either pages constantly (alert fatigue) or gets its threshold raised until real SERVFAIL spikes fall under it.

flowchart TD
  Q[coredns_dns_responses_total by rcode] --> N[NOERROR - answered fine]
  Q --> X[NXDOMAIN - name does not exist]
  Q --> S[SERVFAIL - resolution failed]
  Q --> R[REFUSED - query rejected]
  X --> NX[Expected background in K8s: ndots search-domain expansion]
  NX --> NAL[Do not alert on absolute count]
  S --> SF[Upstream down, API loss, config error]
  SF --> AL1[Alert on SERVFAIL/total ratio vs baseline]
  R --> RF[Missing zone, ACL, max_concurrent]
  RF --> AL2[Alert on any sustained nonzero rate]

Why NXDOMAIN volume is huge in Kubernetes by design

If you run CoreDNS as cluster DNS, your NXDOMAIN count is not a symptom. It is an architectural constant.

Kubernetes pods default to ndots:5 in /etc/resolv.conf. Any name with fewer than 5 dots triggers search-domain expansion before the resolver tries the name as given. A lookup for api.stripe.com (2 dots) from a pod produces this sequence:

  1. api.stripe.com.<namespace>.svc.cluster.local - NXDOMAIN
  2. api.stripe.com.svc.cluster.local - NXDOMAIN
  3. api.stripe.com.cluster.local - NXDOMAIN
  4. api.stripe.com. - the actual intended lookup

Every external lookup generates several intermediate NXDOMAINs. In a typical cluster, NXDOMAIN runs at roughly 20-60% of total responses, shifting with traffic mix: deploy a batch job that resolves many external names and NXDOMAIN volume climbs proportionally, with nothing wrong anywhere.

Two consequences follow:

  • Absolute NXDOMAIN counts mean nothing. They scale with query volume and workload shape. An alert on “NXDOMAIN > N” is an alert on “the cluster is doing DNS.”
  • NXDOMAIN belongs in the denominator, not the numerator. Include it in the total when computing a failure ratio. Excluding it from the denominator inflates the SERVFAIL ratio and makes thresholds meaningless.

Negative caching of these NXDOMAINs is correct behavior, not a leak. coredns_cache_entries{type="denial"} growing steadily is the cache doing its job: absorbing repeated lookups for names that do not exist so they never hit the kubernetes plugin or an upstream. Teams periodically “discover” a big denial cache and treat it as a bug. It is not one.

What SERVFAIL actually tells you

SERVFAIL means CoreDNS could not produce an answer. The usual causes:

CauseWhat it looks likeFirst thing to check
Upstream DNS unreachableSERVFAIL from plugin="forward", fast failure, coredns_forward_healthcheck_broken_total incrementingdig @<upstream_ip> . NS +time=1 +tries=1 from the CoreDNS pod
Kubernetes API unreachableSERVFAIL in the cluster.local zone from plugin="kubernetes", external names still resolvecoredns_kubernetes_rest_client_requests_total{code, method, host} by code label
Corefile misconfigurationSERVFAIL scoped to specific zones, starts after a config changeCorefile contents, coredns_reload_failed_total
Pod served traffic before readySERVFAIL for cluster names right after a rollout, self-resolves in secondsWhether readiness probes use /ready on 8181, not /health on 8080

The plugin label on coredns_dns_responses_total{rcode="SERVFAIL"} is your fastest triage shortcut. plugin="forward" points at upstreams. plugin="kubernetes" points at the API watch path. Use it.

The SERVFAIL caching gotcha

CoreDNS caches SERVFAIL responses for 5 seconds by default (the cache plugin’s servfail TTL). A 1-second upstream blip gets amplified: for 5 seconds after the upstream recovers, clients querying the affected name get SERVFAIL straight from cache. Transient failures look bigger and longer than they were. This is one reason SERVFAIL alerts should be ratio-based over a multi-minute window rather than “any SERVFAIL ever.” Brief, self-resolving blips happen; a sustained ratio above baseline does not.

REFUSED is a failure too

REFUSED means CoreDNS rejected the query outright: no server block matched the zone, an ACL dropped it, or the forward plugin’s max_concurrent limit rejected it (check coredns_forward_max_concurrent_rejects_total). A missing catch-all forward zone makes external lookups come back REFUSED. Any sustained nonzero REFUSED rate warrants a ticket-level alert alongside SERVFAIL.

Quick checks

Read-only commands to see your actual rcode mix right now:

# All responses broken down by rcode and plugin
curl -s http://localhost:9153/metrics | grep '^coredns_dns_responses_total'

# SERVFAIL only, with the plugin label for triage
curl -s http://localhost:9153/metrics | grep 'coredns_dns_responses_total' | grep 'rcode="SERVFAIL"'

# REFUSED only
curl -s http://localhost:9153/metrics | grep 'coredns_dns_responses_total' | grep 'rcode="REFUSED"'

# Total request rate for context (the ratio denominator)
curl -s http://localhost:9153/metrics | grep '^coredns_dns_requests_total'

# Denial cache size - confirm negative caching, don't panic about it
curl -s http://localhost:9153/metrics | grep '^coredns_cache_entries'

In Kubernetes, hit the metrics endpoint on each CoreDNS pod individually (port 9153), or via your existing Prometheus scrape. Replicas behind the kube-dns Service each report their own counters: one degraded replica is invisible in a fleet average.

How to build the alert correctly

  1. Compute the SERVFAIL ratio, not the count. Numerator: rate of coredns_dns_responses_total{rcode="SERVFAIL"} summed across pods. Denominator: rate of all coredns_dns_responses_total. The denominator deliberately includes NXDOMAIN.
  2. Gate on minimum traffic. A 5% SERVFAIL ratio on 0.2 queries per second is one failed query. Require a minimum total query rate (10 qps is a reasonable starting point; tune to your fleet) before the ratio means anything.
  3. Set thresholds against your baseline. A working starting point: above 1% sustained over 5 minutes is degraded and warrants a ticket; above 5% sustained is critical. In a healthy cluster the SERVFAIL ratio sits near zero, so even these thresholds leave room. The important thing is that they trigger on SERVFAIL alone.
  4. Suppress known noise windows. Cold starts after restarts and mass rollouts can produce short SERVFAIL bursts (cache warming, pods serving before ready). Only page when the ratio is corroborated (upstream healthcheck failures, API errors) and the pod has been up long enough that warmup is over.
  5. Alert separately on REFUSED. Any sustained nonzero REFUSED rate is a ticket: it means misconfiguration or forward capacity limits, and it does not fluctuate with traffic the way NXDOMAIN does.
  6. Never alert on absolute NXDOMAIN. If you want NXDOMAIN visibility at all, track the NXDOMAIN/total ratio against its rolling baseline. A ratio change is meaningful (new misconfigured app, reconnaissance, search-domain storm); the absolute count is not.

Metrics and signals to monitor

SignalWhy it mattersWarning sign
coredns_dns_responses_total{rcode="SERVFAIL"}The real availability signal for resolution failuresRatio > 1% of total responses sustained 5m
coredns_dns_responses_total{rcode="SERVFAIL"} by pluginTells you which plugin failed: forward vs kubernetesConcentration in one plugin isolates the cause
coredns_dns_responses_total{rcode="REFUSED"}Policy or capacity rejectionAny sustained nonzero rate
coredns_dns_responses_total{rcode="NXDOMAIN"}Context only: search-domain noise, workload shapeRatio change vs baseline, never absolute count
coredns_forward_healthcheck_broken_totalAll upstreams down: corroborates a SERVFAIL spikeAny increment
coredns_forward_healthcheck_failures_totalPer-upstream healthcheck failures behind SERVFAILsSustained delta for one upstream
coredns_forward_max_concurrent_rejects_totalCapacity-driven REFUSEDAny nonzero rate
coredns_cache_entries{type="denial"}Confirms negative caching is workingDrop to zero = restart or flush; growth = normal

Prevention

  • Delete any alert on combined error counts. If a rule matches rcode!="NOERROR" or sums rcodes into “errors,” rewrite it now. It is either paging on noise or thresholded so high it cannot catch SERVFAIL.
  • Dashboard rcodes as a stacked ratio. NOERROR, NXDOMAIN, SERVFAIL, REFUSED as percentages of total. A healthy cluster shows a stable NXDOMAIN band with SERVFAIL and REFUSED at zero. Deviations are visible at a glance, which makes “why did this alert fire” trivial to answer.
  • Fix NXDOMAIN at the source, not in the alert. If a workload generates extreme search-domain amplification, the right fixes are ndots tuning in the pod spec or fully qualified names with a trailing dot in application config. That reduces noise and upstream load without touching failure detection.
  • Keep per-replica visibility. Aggregate the ratio across pods for the alert, but keep per-pod breakdowns available. One pod SERVFAILing while the other is healthy is a real incident that an average hides.

How Netdata helps

  • Per-rcode response rates out of the box: Netdata charts coredns_dns_responses_total split by rcode and plugin, so NXDOMAIN and SERVFAIL are visually separate streams instead of one blended error counter.
  • Ratio context: plotting SERVFAIL rate next to total request rate makes the SERVFAIL/total ratio legible during an incident without ad-hoc PromQL, and shows when a spike is real versus a low-traffic artifact.
  • Corroboration on one screen: SERVFAIL spikes land next to forward healthcheck failures, per-upstream latency, and Kubernetes API request errors, so the “forward plugin vs kubernetes plugin” triage step takes seconds.
  • Replica-level views: per-pod CoreDNS charts expose single-replica degradation that fleet aggregates mask, which is where quiet SERVFAIL incidents usually live.
  • Cache visibility: success vs denial cache entries sit alongside response codes, so you can confirm a high NXDOMAIN rate is being absorbed by the negative cache rather than hammering upstreams.