The only agent that thinks for itself

Autonomous Monitoring with self-learning AI built-in, operating independently across your entire stack.

Unlimited Metrics & Logs
Machine learning & MCP
5% CPU, 150MB RAM
3GB disk, >1 year retention
800+ integrations, zero config
Dashboards, alerts out of the box
> Discover Netdata Agents

Centralized metrics streaming and storage

Aggregate metrics from multiple agents into centralized Parent nodes for unified monitoring across your infrastructure.

Stream from unlimited agents
Long-term data retention
High availability clustering
Data replication & backup
Scalable architecture
Enterprise-grade security
> Learn about Parents

Fully managed cloud platform

Access your monitoring data from anywhere with our SaaS platform. No infrastructure to manage, automatic updates, and global availability.

Zero infrastructure management
99.9% uptime SLA
Global data centers
Automatic updates & patches
Enterprise SSO & RBAC
SOC2 & ISO certified
> Explore Netdata Cloud

Deploy Netdata Cloud in your infrastructure

Run the full Netdata Cloud platform on-premises for complete data sovereignty and compliance with your security policies.

Complete data sovereignty
Air-gapped deployment
Custom compliance controls
Private network integration
Dedicated support team
Kubernetes & Docker support
> Learn about Cloud On-Premises

Powerful, intuitive monitoring interface

Modern, responsive UI built for real-time troubleshooting with customizable dashboards and advanced visualization capabilities.

Real-time chart updates
Customizable dashboards
Dark & light themes
Advanced filtering & search
Responsive on all devices
Collaboration features
> Explore Netdata UI

Monitor on the go

Native iOS and Android apps bring full monitoring capabilities to your mobile device with real-time alerts and notifications.

iOS & Android apps
Push notifications
Touch-optimized interface
Offline data access
Biometric authentication
Widget support
> Download apps

The future of infrastructure observability

See our strategic direction across AI-native observability, full-stack signals, operational intelligence, and enterprise platform maturity.

AI-native observability
Full-stack signal coverage
Operational intelligence
Enterprise platform maturity
Agent releases every 6 weeks
Cloud continuous delivery
> Explore Product Roadmap

Best energy efficiency

True real-time per-second

100% automated zero config

Centralized observability

Multi-year retention

High availability built-in

Zero maintenance

Always up-to-date

Enterprise security

Complete data control

Air-gap ready

Compliance certified

Millisecond responsiveness

Infinite zoom & pan

Works on any device

Native performance

Instant alerts

Monitor anywhere

AI-native observability

Continuous delivery

Open source foundation

80% Faster Incident Resolution

AI-powered troubleshooting from detection, to root cause and blast radius identification, to reporting.

True Real-Time and Simple, even at Scale

Linearly and infinitely scalable full-stack observability, that can be deployed even mid-crisis.

90% Cost Reduction, Full Fidelity

Instead of centralizing the data, Netdata distributes the code, eliminating pipelines and complexity.

See and Map Your Entire Network

Live topology, flow analytics, and SNMP device and trap monitoring — unified with your full-stack observability.

Control Without Surrender

SOC 2 Type 2 certified with every metric kept on your infrastructure.

Integrations

800+ collectors and notification channels, auto-discovered and ready out of the box.

800+ data collectors
Auto-discovery & zero config
Cloud, infra, app protocols
Notifications out of the box
> Explore integrations
Real Results
46% Cost Reduction

Reduced monitoring costs by 46% while cutting staff overhead by 67%.

— Leonardo Antunez, Codyas

Zero Pipeline

No data shipping. No central storage costs. Query at the edge.

From Our Users
"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

No Query Language

Point-and-click troubleshooting. No PromQL, no LogQL, no learning curve.

Enterprise Ready
67% Less Staff, 46% Cost Cut

Enterprise efficiency without enterprise complexity—real ROI from day one.

— Leonardo Antunez, Codyas

SOC 2 Type 2 Certified

Zero data egress. Only metadata reaches the cloud. Your metrics stay on your infrastructure.

Full Coverage
800+ Collectors

Auto-discovered and configured. No manual setup required.

Any Notification Channel

Slack, PagerDuty, Teams, email, webhooks—all built-in.

Built for the People Who Get Paged

Because 3am alerts deserve instant answers, not hour-long hunts.

Every Industry Has Rules. We Master Them.

See how healthcare, finance, and government teams cut monitoring costs 90% while staying audit-ready.

Monitor Any Technology. Configure Nothing.

Install the agent. It already knows your stack.
From Our Users
"A Rare Unicorn"

Netdata gives more than you invest in it. A rare unicorn that obeys the Pareto rule.

— Eduard Porquet Mateu, TMB Barcelona

99% Downtime Reduction

Reduced website downtime by 99% and cloud bill by 30% using Netdata alerts.

— Falkland Islands Government

Real Savings
30% Cloud Cost Reduction

Optimized resource allocation based on Netdata alerts cut cloud spending by 30%.

— Falkland Islands Government

46% Cost Cut

Reduced monitoring staff by 67% while cutting operational costs by 46%.

— Codyas

Real Coverage
"Plugin for Everything"

Netdata has agent capacity or a plugin for everything, including Windows and Kubernetes.

— Eduard Porquet Mateu, TMB Barcelona

"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

Real Speed
Troubleshooting in 30 Seconds

From 2-3 minutes to 30 seconds—instant visibility into any node issue.

— Matthew Artist, Nodecraft

20% Downtime Reduction

20% less downtime and 40% budget optimization from out-of-the-box monitoring.

— Simon Beginn, LANCOM Systems

Pay per Node. Unlimited Everything Else.

One price per node. Unlimited metrics, logs, users, and retention. No per-GB surprises.

Free tier—forever
No metric limits or caps
Retention you control
Cancel anytime
> See pricing plans

What's Your Monitoring Really Costing You?

Most teams overpay by 40-60%. Let's find out why.

Expose hidden metric charges
Calculate tool consolidation
Customers report 30-67% savings
Results in under 60 seconds
> See what you're really paying

Your Infrastructure Is Unique. Let's Talk.

Because monitoring 10 nodes is different from monitoring 10,000.

On-prem & air-gapped deployment
Volume pricing & agreements
Architecture review for your scale
Compliance & security support
> Start a conversation

Monitoring That Sells Itself

Deploy in minutes. Impress clients in hours. Earn recurring revenue for years.

30-second live demos close deals
Zero config = zero support burden
Competitive margins & deal protection
Response in 48 hours
> Apply to partner

Per-Second Metrics at Homelab Prices

Same engine, same dashboards, same ML. Just priced for tinkerers.

Community: Free forever · 5 nodes · non-commercial
Homelab: $90/yr · unlimited nodes · fair usage
> Get the Homelab Plan

$1,000 Per Referral. Unlimited Referrals.

Your colleagues get 10% off. You get 10% commission. Everyone wins.

10% of subscriptions, up to $1,000 each
Track earnings inside Netdata Cloud
PayPal/Venmo payouts in 3-4 weeks
No caps, no complexity
> Get your referral link
Cost Proof
40% Budget Optimization

"Netdata's significant positive impact" — LANCOM Systems

Calculate Your Savings

Compare vs Datadog, Grafana, Dynatrace

Savings Proof
46% Cost Reduction

"Cut costs by 46%, staff by 67%" — Codyas

30% Cloud Bill Savings

"Reduced cloud bill by 30%" — Falkland Islands Gov

Enterprise Proof
"Better Than Combined Alternatives"

"Better observability with Netdata than combining other tools." — TMB Barcelona

Real Engineers, <24h Response

DPA, SLAs, on-prem, volume pricing

Why Partners Win
Demo Live Infrastructure

One command, 30 seconds, real data—no sandbox needed

Zero Tickets, High Margins

Auto-config + per-node pricing = predictable profit

Homelab Ready
Free Video Course

8-episode Netdata tutorial by LearnLinux.tv

76k+ GitHub Stars

3rd most starred monitoring project

Worth Recommending
Product That Delivers

Customers report 40-67% cost cuts, 99% downtime reduction

Zero Risk to Your Rep

Free tier lets them try before they buy

AI Support Assistant, Available 24/7

Nedi has access to all official documentation, source code, and resources. Ask any question about Netdata—responds in your language.

Deployment & configuration
Troubleshooting & sizing
Alerts & notifications
Evidence-based answers
> Ask Nedi now

Never Fight Fires Alone

Docs, community, and expert help—pick your path to resolution.

Learn.netdata.cloud docs
Discord, Forums, GitHub
Premium support available
> Get answers now

60 Seconds to First Dashboard

One command to install. Zero config. 850+ integrations documented.

Linux, Windows, K8s, Docker
Auto-discovers your stack
> Read our documentation

76,000+ Engineers Strong

615+ contributors. 1.5M daily downloads. One mission: simplify observability.

Per-Second. 90% Cheaper. Data Stays Home.

Side-by-side comparisons: costs, real-time granularity, and data sovereignty for every major tool.

See why teams switch from Datadog, Prometheus, Grafana, and more.

> Browse all comparisons
Edge-Native Observability, Born Open Source
Per-second visibility, ML on every metric, and data that never leaves your infrastructure.
Founded in 2016
615+ contributors worldwide
Remote-first, engineering-driven
Open source first
> Read our story
Promises We Publish—and Prove
12 principles backed by open code, independent validation, and measurable outcomes.
Open source, peer-reviewed
Zero config, instant value
Data sovereignty by design
Aligned pricing, no surprises
> See all 12 principles
Edge-Native, AI-Ready, 100% Open
76k+ stars. Full ML, AI, and automation—GPLv3+, not premium add-ons.
76,000+ GitHub stars
GPLv3+ licensed forever
ML on every metric, included
Zero vendor lock-in
> Explore our open source
Build Real-Time Observability for the World
Remote-first team shipping per-second monitoring with ML on every metric.
Remote-first, fully distributed
Open source (76k+ stars)
Challenging technical problems
Your code on millions of systems
> See open roles
Meet the Team Behind Netdata
Conferences, meetups, and tradeshows where you can see Netdata in action and talk to the engineers who build it.
Live demos and deep dives
Book 1-on-1 meetings
Talks and panel sessions
Event recaps and photos
> See all events
Talk to a Netdata Human in <24 Hours
Sales, partnerships, press, or professional services—real engineers, fast answers.
Discuss your observability needs
Pricing and volume discounts
Partnership opportunities
Media and press inquiries
> Book a conversation
Your Data. Your Rules.
On-prem data, cloud control plane, transparent terms.
Trust & Scale
76,000+ GitHub Stars

One of the most popular open-source monitoring projects

SOC 2 Type 2 Certified

Enterprise-grade security and compliance

Data Sovereignty

Your metrics stay on your infrastructure

Validated
University of Amsterdam

"Most energy-efficient monitoring solution" — ICSOC 2023, peer-reviewed

ADASTEC (Autonomous Driving)

"Doesn't miss alerts—mission-critical trust for safety software"

Community Stats
615+ Contributors

Global community improving monitoring for everyone

1.5M+ Downloads/Day

Trusted by teams worldwide

GPLv3+ Licensed

Free forever, fully open source agent

Why Join?
Remote-First

Work from anywhere, async-friendly culture

Impact at Scale

Your work helps millions of systems

$ guides / coredns / coredns-nodelocaldns-monitoring ▌

Operations Guides

CoreDNS with NodeLocal DNSCache: what changes about monitoring

You deployed NodeLocal DNSCache to fix the 5-second timeout race and conntrack pressure, and it worked. DNS latency dropped, the conntrack table stopped filling, and CoreDNS QPS fell off a cliff. Then someone looked at the CoreDNS dashboard, saw near-zero traffic, and concluded CoreDNS was oversized. That conclusion is how the next incident starts.

NodeLocal DNSCache runs a caching DNS proxy as a DaemonSet on every node. It intercepts pod DNS queries before they enter the iptables DNAT and conntrack path, so each query is answered on the local node instead of traversing NAT to a central CoreDNS pod. This eliminates the kernel race condition behind the 5-second glibc timeout and removes DNS as a conntrack consumer.

The cost is that your monitoring model inverts. Most queries never reach central CoreDNS anymore. CoreDNS QPS now measures node-local cache misses, not cluster DNS demand. Every dashboard, alert threshold, and capacity assumption built on “CoreDNS traffic equals DNS traffic” is now wrong.

What changes when you deploy NodeLocal DNSCache

Without NodeLocal DNSCache, the query path is: pod to kube-dns ClusterIP, through iptables DNAT and conntrack, to a CoreDNS pod. CoreDNS sees every query. Its metrics are a complete picture of cluster DNS demand.

With NodeLocal DNSCache, each node runs a local caching instance (itself CoreDNS, configured as a caching forwarder). Pods are pointed at a link-local address on the node, typically 169.254.20.10. The node-local instance answers from its cache when it can. On a miss, it forwards to the central CoreDNS service, and the node-local Corefile forces TCP for that upstream hop by default, which changes the connection profile between the layers.

Two consequences follow:

  • The central CoreDNS becomes a miss handler. Its traffic is whatever the per-node caches could not serve: cold entries, expired TTLs, names with short or zero TTLs, and the long tail of unique names. This traffic is bursty and correlates across nodes during events like rollouts or upstream TTL expiry.
  • Failure blast radius moves. A central CoreDNS or upstream failure no longer degrades gracefully per query. Node-local caches absorb it silently until their entries expire, then every node re-queries at once. The failure arrives at central CoreDNS as a synchronized flood.
flowchart LR
  subgraph node["Every node"]
    pod["Application pod"] -->|"query to link-local IP"| nld["node-local-dns cache"]
    nld -->|"cache hit"| pod
  end
  nld -->|"miss, TCP"| cd["Central CoreDNS"]
  cd -->|"cluster.local"| k8s["Kubernetes API state"]
  cd -->|"external zones"| up["Upstream resolvers"]

The old path, pod to ClusterIP through NAT, is gone from the hot path. That is exactly why the signals you used to rely on, conntrack pressure and CoreDNS-side query volume, no longer tell you about total demand.

What changes on the central CoreDNS side

QPS no longer means load

coredns_dns_requests_total on the central CoreDNS pods now counts cache misses from every node-local instance. A drop of 80 to 95 percent after enabling NodeLocal DNSCache is expected and healthy. It is not spare capacity you can reclaim.

The trap is under-provisioning central CoreDNS because “it was not doing much.” When an upstream resolver fails or a wave of TTLs expires, every node’s cache misses simultaneously and the flood lands on a CoreDNS deployment that was scaled down to match its quiet period. Size central CoreDNS for the miss storm, not the steady state.

Latency gets slower on average, and that is fine

Central CoreDNS used to serve a mix of cache hits (sub-millisecond) and forwards. Now node-local caches absorb the hits, so coredns_dns_request_duration_seconds on central CoreDNS skews toward forwarded and kubernetes-plugin queries. Average latency rises. P99 matters more than the mean here; watch the forward plugin’s per-upstream latency (coredns_proxy_request_duration_seconds{proxy_name="forward"} on CoreDNS 1.11.0+, coredns_forward_request_duration_seconds on 1.5 through 1.10) to see whether the rise is your upstreams or your CoreDNS.

Cache metrics at the central layer change meaning

The central cache now mostly holds entries that node-local caches also missed on. Its hit ratio (coredns_cache_hits_total / coredns_cache_requests_total) will typically fall, because the easy, repetitive hits are served one layer down. Do not alert on the central cache hit ratio dropping after the rollout; establish a new baseline first. Evictions (coredns_cache_evictions_total) still mean the central cache is undersized for its working set, but the working set itself changed.

SERVFAIL amplification works differently now

CoreDNS caches SERVFAIL for 5 seconds by default, amplifying brief upstream blips. With NodeLocal DNSCache there are two caching layers. A short upstream failure can be cached at the central layer and then re-cached by every node-local instance, so a one-second blip can surface as cluster-wide failures that outlive the actual upstream problem. Watch coredns_dns_responses_total{rcode="SERVFAIL"} at both layers, split by the plugin label, before assuming the upstream is still down.

What appears on the node-local-dns side

The node-local-dns pods run CoreDNS in caching mode and expose the same Prometheus metric families, but on port 9253 rather than 9153. Because the DaemonSet runs one pod per node, every signal is per-node. There is no cluster-wide aggregate unless you build one by summing across pods.

What the layer tells you:

  • Real cluster DNS demand. coredns_dns_requests_total summed across all node-local-dns pods is the true cluster QPS. This is the number central CoreDNS QPS used to represent.
  • Node-local cache effectiveness. The hit ratio per node shows how much traffic never leaves the node. Low hit ratio on one node with normal ratios elsewhere points at a node-specific workload pattern, not a DNS problem.
  • Miss pressure heading for central CoreDNS. Miss rate across all nodes, aggregated, is the best leading indicator of load about to hit the central layer.
  • Upstream health from the edge. The node-local instances health-check their upstream, which is central CoreDNS. A simultaneous spike in health-check failures across many nodes is your earliest warning that central CoreDNS is struggling, before SERVFAILs reach applications.
  • Connection pressure. The forced-TCP hop to central CoreDNS means each node-local instance holds upstream connections. Watch process_open_fds and the connection cache metrics (coredns_proxy_conn_cache_misses_total{proxy_name="forward"} on CoreDNS 1.11.0+, coredns_forward_conn_cache_misses_total on 1.5 through 1.10) on busy nodes.

One operational caveat: the node-local-dns pod installs iptables rules to intercept DNS traffic. If the pod is OOMKilled, those rules can persist while the cache is down, producing a DNS gap on that node until the container restarts. Monitor per-node pod restarts, not just aggregate DaemonSet health.

The failure mode this architecture hides

The most dangerous property of the two-layer design is how well it masks central-layer failure. A warm cache hides an upstream failure until TTLs expire, then the failure becomes sudden and complete. NodeLocal DNSCache multiplies the effect because the masking cache is distributed across every node and the unmasking is synchronized.

The sequence looks like this:

  1. Central CoreDNS or its upstreams degrade. Node-local caches keep answering from warm entries. Application metrics stay green.
  2. TTLs expire. Entries age out at roughly the same rate across nodes because they were cached at roughly the same time.
  3. Every node-local instance starts forwarding. Central CoreDNS goes from quiet to flooded in seconds. If it was scaled down during the quiet period, it now hits coredns_forward_max_concurrent_rejects_total, REFUSED responses, and goroutine accumulation.
  4. Node-local caches receive and re-cache failures, extending the blast radius beyond the original fault.

Your early-warning signals for this sequence all live at the seams: node-local health-check failures toward central CoreDNS, central coredns_forward_healthcheck_broken_total and per-upstream coredns_proxy_healthcheck_failures_total{proxy_name="forward"}, and the aggregate node-local miss rate. End-to-end success ratio will be the last signal to move, which is exactly why you cannot rely on it.

Signals to watch in production

SignalLayerWhy it mattersWarning sign
coredns_dns_requests_total summed over node-local-dns podsNode-localTrue cluster DNS demandDrift from established baseline; nobody else measures this
Node-local cache hit ratioNode-localHow much load central CoreDNS is shielded fromSustained drop means more miss traffic heading to the central layer
Node-local upstream health-check failuresNode-localEarliest view of central CoreDNS troubleSimultaneous rise across many nodes
coredns_dns_requests_total on central CoreDNSCentralMiss-handling load onlyUsing it for capacity decisions; sudden spikes mean a synchronized miss storm
coredns_dns_responses_total{rcode="SERVFAIL"} by pluginBothActual user pain, attributed to a layerAny sustained nonzero rate; divergence between layers indicates cache re-poisoning
coredns_proxy_healthcheck_failures_total{proxy_name="forward"} per upstream (to)CentralWhich upstream is failing before all failSustained increments for one upstream
coredns_forward_healthcheck_broken_totalCentralAll upstreams unhealthyAny increment; with warm node-local caches this fires before users notice
go_goroutines and go_memstats_heap_inuse_bytesCentralAccumulation during a miss floodGrowth that does not return to baseline after the event
process_open_fdsNode-localTCP connections to central CoreDNS on busy nodesSteady climb toward process_max_fds

Two habits make this table usable. First, always split dashboards by layer; a chart that mixes central and node-local CoreDNS metrics with the same metric names will mislead you. Second, baseline everything again after the rollout. Pre-NodeLocal thresholds for QPS, hit ratio, and latency are invalid by construction.

Common misreadings after the rollout

  • “CoreDNS QPS dropped 90 percent, something is broken.” It is working. Verify by checking that aggregate node-local QPS matches your pre-rollout cluster demand.
  • “CoreDNS is idle, scale it down.” It is a miss handler. Scale it for the correlated miss storm during an upstream failure or mass TTL expiry, and keep enough replicas that a rolling restart does not coincide with one.
  • “DNS looks healthy, no alerts fired.” Check whether your alerts only watched the central layer. Node-local failures, per-node OOM kills, and rising miss rates are invisible there.
  • “Latency went up after we deployed the cache.” Central-layer latency rising is expected because the hits moved. Compare application-observed DNS latency before and after, not the central CoreDNS histogram.

How Netdata helps

  • Per-second scraping of both layers makes the synchronized miss flood visible as it forms, where minute-resolution monitoring shows only the aftermath.
  • Netdata discovers CoreDNS Prometheus endpoints and charts the standard metric families, so you can place node-local-dns (port 9253) and central CoreDNS (port 9153) side by side and compare QPS, SERVFAIL by rcode, and cache hit ratios on the same time axis.
  • Per-upstream breakdowns of request duration and health-check failure metrics on central CoreDNS let you see a single failing upstream while node-local caches are still masking it from applications.
  • Go runtime metrics (go_goroutines, heap, GC pauses) on central CoreDNS expose the accumulation phase of a miss storm before pods start rejecting queries or getting OOMKilled.
  • Node-level views alongside DaemonSet pod metrics help you spot the single-node cases: an OOMKilled node-local-dns pod, FD exhaustion on a busy node, or a workload pattern producing an abnormally low local hit ratio.
  • ML-based anomaly detection on the aggregate node-local miss rate gives you a leading indicator for central-layer load without hand-tuned thresholds on a traffic pattern that changes with every rollout.