The only agent that thinks for itself

Autonomous Monitoring with self-learning AI built-in, operating independently across your entire stack.

Unlimited Metrics & Logs
Machine learning & MCP
5% CPU, 150MB RAM
3GB disk, >1 year retention
800+ integrations, zero config
Dashboards, alerts out of the box
> Discover Netdata Agents

Centralized metrics streaming and storage

Aggregate metrics from multiple agents into centralized Parent nodes for unified monitoring across your infrastructure.

Stream from unlimited agents
Long-term data retention
High availability clustering
Data replication & backup
Scalable architecture
Enterprise-grade security
> Learn about Parents

Fully managed cloud platform

Access your monitoring data from anywhere with our SaaS platform. No infrastructure to manage, automatic updates, and global availability.

Zero infrastructure management
99.9% uptime SLA
Global data centers
Automatic updates & patches
Enterprise SSO & RBAC
SOC2 & ISO certified
> Explore Netdata Cloud

Deploy Netdata Cloud in your infrastructure

Run the full Netdata Cloud platform on-premises for complete data sovereignty and compliance with your security policies.

Complete data sovereignty
Air-gapped deployment
Custom compliance controls
Private network integration
Dedicated support team
Kubernetes & Docker support
> Learn about Cloud On-Premises

Powerful, intuitive monitoring interface

Modern, responsive UI built for real-time troubleshooting with customizable dashboards and advanced visualization capabilities.

Real-time chart updates
Customizable dashboards
Dark & light themes
Advanced filtering & search
Responsive on all devices
Collaboration features
> Explore Netdata UI

Monitor on the go

Native iOS and Android apps bring full monitoring capabilities to your mobile device with real-time alerts and notifications.

iOS & Android apps
Push notifications
Touch-optimized interface
Offline data access
Biometric authentication
Widget support
> Download apps

The future of infrastructure observability

See our strategic direction across AI-native observability, full-stack signals, operational intelligence, and enterprise platform maturity.

AI-native observability
Full-stack signal coverage
Operational intelligence
Enterprise platform maturity
Agent releases every 6 weeks
Cloud continuous delivery
> Explore Product Roadmap

Best energy efficiency

True real-time per-second

100% automated zero config

Centralized observability

Multi-year retention

High availability built-in

Zero maintenance

Always up-to-date

Enterprise security

Complete data control

Air-gap ready

Compliance certified

Millisecond responsiveness

Infinite zoom & pan

Works on any device

Native performance

Instant alerts

Monitor anywhere

AI-native observability

Continuous delivery

Open source foundation

80% Faster Incident Resolution

AI-powered troubleshooting from detection, to root cause and blast radius identification, to reporting.

True Real-Time and Simple, even at Scale

Linearly and infinitely scalable full-stack observability, that can be deployed even mid-crisis.

90% Cost Reduction, Full Fidelity

Instead of centralizing the data, Netdata distributes the code, eliminating pipelines and complexity.

See and Map Your Entire Network

Live topology, flow analytics, and SNMP device and trap monitoring — unified with your full-stack observability.

Control Without Surrender

SOC 2 Type 2 certified with every metric kept on your infrastructure.

Integrations

800+ collectors and notification channels, auto-discovered and ready out of the box.

800+ data collectors
Auto-discovery & zero config
Cloud, infra, app protocols
Notifications out of the box
> Explore integrations
Real Results
46% Cost Reduction

Reduced monitoring costs by 46% while cutting staff overhead by 67%.

— Leonardo Antunez, Codyas

Zero Pipeline

No data shipping. No central storage costs. Query at the edge.

From Our Users
"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

No Query Language

Point-and-click troubleshooting. No PromQL, no LogQL, no learning curve.

Enterprise Ready
67% Less Staff, 46% Cost Cut

Enterprise efficiency without enterprise complexity—real ROI from day one.

— Leonardo Antunez, Codyas

SOC 2 Type 2 Certified

Zero data egress. Only metadata reaches the cloud. Your metrics stay on your infrastructure.

Full Coverage
800+ Collectors

Auto-discovered and configured. No manual setup required.

Any Notification Channel

Slack, PagerDuty, Teams, email, webhooks—all built-in.

Built for the People Who Get Paged

Because 3am alerts deserve instant answers, not hour-long hunts.

Every Industry Has Rules. We Master Them.

See how healthcare, finance, and government teams cut monitoring costs 90% while staying audit-ready.

Monitor Any Technology. Configure Nothing.

Install the agent. It already knows your stack.
From Our Users
"A Rare Unicorn"

Netdata gives more than you invest in it. A rare unicorn that obeys the Pareto rule.

— Eduard Porquet Mateu, TMB Barcelona

99% Downtime Reduction

Reduced website downtime by 99% and cloud bill by 30% using Netdata alerts.

— Falkland Islands Government

Real Savings
30% Cloud Cost Reduction

Optimized resource allocation based on Netdata alerts cut cloud spending by 30%.

— Falkland Islands Government

46% Cost Cut

Reduced monitoring staff by 67% while cutting operational costs by 46%.

— Codyas

Real Coverage
"Plugin for Everything"

Netdata has agent capacity or a plugin for everything, including Windows and Kubernetes.

— Eduard Porquet Mateu, TMB Barcelona

"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

Real Speed
Troubleshooting in 30 Seconds

From 2-3 minutes to 30 seconds—instant visibility into any node issue.

— Matthew Artist, Nodecraft

20% Downtime Reduction

20% less downtime and 40% budget optimization from out-of-the-box monitoring.

— Simon Beginn, LANCOM Systems

Pay per Node. Unlimited Everything Else.

One price per node. Unlimited metrics, logs, users, and retention. No per-GB surprises.

Free tier—forever
No metric limits or caps
Retention you control
Cancel anytime
> See pricing plans

What's Your Monitoring Really Costing You?

Most teams overpay by 40-60%. Let's find out why.

Expose hidden metric charges
Calculate tool consolidation
Customers report 30-67% savings
Results in under 60 seconds
> See what you're really paying

Your Infrastructure Is Unique. Let's Talk.

Because monitoring 10 nodes is different from monitoring 10,000.

On-prem & air-gapped deployment
Volume pricing & agreements
Architecture review for your scale
Compliance & security support
> Start a conversation

Monitoring That Sells Itself

Deploy in minutes. Impress clients in hours. Earn recurring revenue for years.

30-second live demos close deals
Zero config = zero support burden
Competitive margins & deal protection
Response in 48 hours
> Apply to partner

Per-Second Metrics at Homelab Prices

Same engine, same dashboards, same ML. Just priced for tinkerers.

Community: Free forever · 5 nodes · non-commercial
Homelab: $90/yr · unlimited nodes · fair usage
> Get the Homelab Plan

$1,000 Per Referral. Unlimited Referrals.

Your colleagues get 10% off. You get 10% commission. Everyone wins.

10% of subscriptions, up to $1,000 each
Track earnings inside Netdata Cloud
PayPal/Venmo payouts in 3-4 weeks
No caps, no complexity
> Get your referral link
Cost Proof
40% Budget Optimization

"Netdata's significant positive impact" — LANCOM Systems

Calculate Your Savings

Compare vs Datadog, Grafana, Dynatrace

Savings Proof
46% Cost Reduction

"Cut costs by 46%, staff by 67%" — Codyas

30% Cloud Bill Savings

"Reduced cloud bill by 30%" — Falkland Islands Gov

Enterprise Proof
"Better Than Combined Alternatives"

"Better observability with Netdata than combining other tools." — TMB Barcelona

Real Engineers, <24h Response

DPA, SLAs, on-prem, volume pricing

Why Partners Win
Demo Live Infrastructure

One command, 30 seconds, real data—no sandbox needed

Zero Tickets, High Margins

Auto-config + per-node pricing = predictable profit

Homelab Ready
Free Video Course

8-episode Netdata tutorial by LearnLinux.tv

76k+ GitHub Stars

3rd most starred monitoring project

Worth Recommending
Product That Delivers

Customers report 40-67% cost cuts, 99% downtime reduction

Zero Risk to Your Rep

Free tier lets them try before they buy

AI Support Assistant, Available 24/7

Nedi has access to all official documentation, source code, and resources. Ask any question about Netdata—responds in your language.

Deployment & configuration
Troubleshooting & sizing
Alerts & notifications
Evidence-based answers
> Ask Nedi now

Never Fight Fires Alone

Docs, community, and expert help—pick your path to resolution.

Learn.netdata.cloud docs
Discord, Forums, GitHub
Premium support available
> Get answers now

60 Seconds to First Dashboard

One command to install. Zero config. 850+ integrations documented.

Linux, Windows, K8s, Docker
Auto-discovers your stack
> Read our documentation

76,000+ Engineers Strong

615+ contributors. 1.5M daily downloads. One mission: simplify observability.

Per-Second. 90% Cheaper. Data Stays Home.

Side-by-side comparisons: costs, real-time granularity, and data sovereignty for every major tool.

See why teams switch from Datadog, Prometheus, Grafana, and more.

> Browse all comparisons
Edge-Native Observability, Born Open Source
Per-second visibility, ML on every metric, and data that never leaves your infrastructure.
Founded in 2016
615+ contributors worldwide
Remote-first, engineering-driven
Open source first
> Read our story
Promises We Publish—and Prove
12 principles backed by open code, independent validation, and measurable outcomes.
Open source, peer-reviewed
Zero config, instant value
Data sovereignty by design
Aligned pricing, no surprises
> See all 12 principles
Edge-Native, AI-Ready, 100% Open
76k+ stars. Full ML, AI, and automation—GPLv3+, not premium add-ons.
76,000+ GitHub stars
GPLv3+ licensed forever
ML on every metric, included
Zero vendor lock-in
> Explore our open source
Build Real-Time Observability for the World
Remote-first team shipping per-second monitoring with ML on every metric.
Remote-first, fully distributed
Open source (76k+ stars)
Challenging technical problems
Your code on millions of systems
> See open roles
Meet the Team Behind Netdata
Conferences, meetups, and tradeshows where you can see Netdata in action and talk to the engineers who build it.
Live demos and deep dives
Book 1-on-1 meetings
Talks and panel sessions
Event recaps and photos
> See all events
Talk to a Netdata Human in <24 Hours
Sales, partnerships, press, or professional services—real engineers, fast answers.
Discuss your observability needs
Pricing and volume discounts
Partnership opportunities
Media and press inquiries
> Book a conversation
Your Data. Your Rules.
On-prem data, cloud control plane, transparent terms.
Trust & Scale
76,000+ GitHub Stars

One of the most popular open-source monitoring projects

SOC 2 Type 2 Certified

Enterprise-grade security and compliance

Data Sovereignty

Your metrics stay on your infrastructure

Validated
University of Amsterdam

"Most energy-efficient monitoring solution" — ICSOC 2023, peer-reviewed

ADASTEC (Autonomous Driving)

"Doesn't miss alerts—mission-critical trust for safety software"

Community Stats
615+ Contributors

Global community improving monitoring for everyone

1.5M+ Downloads/Day

Trusted by teams worldwide

GPLv3+ Licensed

Free forever, fully open source agent

Why Join?
Remote-First

Work from anywhere, async-friendly culture

Impact at Scale

Your work helps millions of systems

$ guides / coredns / coredns-oomkilled ▌

Operations Guides

CoreDNS OOMKilled: the memory cliff and the restart crash loop

You found the CoreDNS pods in kube-system with OOMKilled as the last terminated state. Maybe one pod restarted once and recovered. Maybe both replicas are flapping and cluster DNS is degraded or down. Either way, the pod status operators paste into a search bar is the same: Last State: Terminated, Reason: OOMKilled.

Memory exhaustion in CoreDNS is a cliff edge, not a slowdown. There is no degradation curve where latency rises and errors climb while you watch. RSS approaches the container memory limit, the kernel OOM killer terminates the process instantly, and every client served by that pod loses DNS until a replacement is ready. The warning phase exists only in your metrics, and only if you are watching the right ones.

The nastier variant is the restart crash loop. The pod is OOMKilled, Kubernetes restarts it, the kubernetes plugin does a full re-list of all Services, Endpoints, and Pods to rebuild its in-memory snapshot, and that startup re-list allocates 2-3x steady-state memory. If the limit was set just above steady state, the re-list peak blows straight through it and the pod is OOMKilled again before it ever becomes ready. This is not hypothetical: in the upstream 5,000-node scalability test (kubernetes/kubernetes#139117), CoreDNS at roughly 82MB steady state was consistently OOMKilled by the re-list after an apiserver restart at the default 170MB limit, and the fix was doubling the limit to 340MB. If your limit is tuned to steady state with 10-20% headroom, one restart can put you in the same loop.

What this means

The OOM killer does not watch Go heap. It watches RSS, which is process_resident_memory_bytes on the metrics endpoint. RSS is always larger than live heap (go_memstats_alloc_bytes) because the Go runtime holds freed memory instead of eagerly returning it to the OS. You can have a reasonable heap and still get killed because RSS, GC headroom, goroutine stacks, and runtime overhead pushed past the cgroup limit.

Four things consume CoreDNS memory in practice:

  • The kubernetes plugin’s in-memory snapshot of Services, Endpoints, and optionally Pods. This scales with cluster size and is the dominant consumer in large clusters.
  • The cache plugin’s positive and negative caches, proportional to configured size and average record size.
  • Goroutines. One per in-flight query plus background workers. Slow upstreams cause blocked goroutines to accumulate, and each one holds stack and per-request allocations.
  • Go runtime overhead: GC metadata, freed-but-unreturned memory, allocator behavior.

The crash loop mechanism:

flowchart TD
  A[Steady state RSS near limit] --> B[OOM kill: instant termination]
  B --> C[Pod restarts]
  C --> D[Kubernetes plugin full re-list]
  D --> E[Re-list peak: 2-3x steady state]
  E --> F{Peak under limit?}
  F -->|No| B
  F -->|Yes| G[Ready, serves DNS]
  G --> H[Growth: cluster size, cache, leaks]
  H --> A

The loop breaks only if the re-list peak fits under the limit, or if whatever grew memory past steady state is fixed.

Common causes

CauseWhat it looks likeFirst thing to check
Limit sized for steady state, not the re-list peakOOMKilled immediately after every restart; pod never reaches ready; restarts correlate with apiserver restarts or rolloutsRestart timing vs. OOMKilled events; cluster Service/Endpoint count
Cluster growth outpaced the limitSlow RSS climb over weeks, then first OOMKill; more frequent as Services/Endpoints growkubectl get svc --all-namespaces | wc -l trend vs. limit
Cache too large for the limitCache entries pinned at configured max, evictions high, RSS tracks cache sizecoredns_cache_entries vs. Corefile cache size
Goroutine or connection leakgo_goroutines grows without returning to baseline; RSS climbs in step; not tied to QPSgo_goroutines trend over hours
Blocked goroutines from slow upstreamsGoroutine and memory growth correlates with elevated upstream latencycoredns_proxy_request_duration_seconds per upstream
Unbounded forward concurrencyMemory spikes under query floods; no max_concurrent set in CorefileCorefile forward block

Quick checks

All read-only and safe to run during an incident.

# Pod status and restart counts
kubectl get pods -n kube-system -l k8s-app=kube-dns

# Confirm OOMKilled and see exit details
kubectl describe pod -n kube-system <coredns-pod> | grep -A 10 "Last State"

# Recent OOM events
kubectl get events -n kube-system | grep -i oom

# Previous container logs (what happened right before the kill)
kubectl logs -n kube-system <coredns-pod> --previous --tail=100

The metrics below are on the CoreDNS metrics endpoint (:9153, when the prometheus plugin is enabled). Run them from wherever you normally scrape, for example by port-forwarding (kubectl port-forward -n kube-system <coredns-pod> 9153) and curling from your workstation. The CoreDNS scratch image has no shell, curl, or wget, so kubectl exec into the CoreDNS container cannot fetch metrics.

# Current RSS: this is the number the OOM killer watches
curl -s http://localhost:9153/metrics | grep '^process_resident_memory_bytes'

# Goroutine count: leaks show unbounded growth away from baseline
curl -s http://localhost:9153/metrics | grep '^go_goroutines'

# Heap in use: for leak trend analysis, not OOM prediction
curl -s http://localhost:9153/metrics | grep '^go_memstats_heap_inuse_bytes'

# Cache occupancy vs configured maximum
curl -s http://localhost:9153/metrics | grep '^coredns_cache_entries'
curl -s http://localhost:9153/metrics | grep 'coredns_cache_evictions_total'
# Cluster size driving the kubernetes plugin snapshot
kubectl get svc --all-namespaces | wc -l
kubectl get endpoints --all-namespaces | wc -l

If you have metrics-server installed, kubectl top pod -n kube-system gives a quick current-usage reading, but it is a point sample. The trend before the kill matters more than the value after.

How to diagnose it

  1. Confirm the kill reason. kubectl describe pod should show Reason: OOMKilled under Last State. If it shows something else (Error, Completed, loop detection), this is a different failure. CrashLoopBackOff with “Loop detected” in logs is a forwarding loop, not memory.

  2. Distinguish the crash loop from a slow leak. Look at restart timing. If the pod is killed within seconds to a minute of starting, every time, you are hitting the re-list peak. If the pod runs for hours or days and memory climbs until the kill, you have growth: cluster expansion, cache pressure, or a leak.

  3. Reconstruct the pre-kill memory trend. You need historical process_resident_memory_bytes for the killed pod. Post-GC minimums rising steadily means a leak or an undersized limit for the working set. A flat trend that spikes only at startup means the re-list peak is the whole problem.

  4. Check goroutines. If go_goroutines climbed in step with RSS and never returned to baseline after load dropped, suspect a leak or blocked upstream calls. Correlate with per-upstream latency (coredns_proxy_request_duration_seconds{to=...}): slow-but-not-dead upstreams accumulate blocked goroutines, and each one holds memory.

  5. Check the cache. Compare coredns_cache_entries against the size in your Corefile cache directive. Entries pinned at the maximum with a rising eviction rate means the working set exceeds the cache, but the cache itself is also consuming the heap it was allocated. The default cache is 9984 items per cache (success and denial are separate), roughly 30MB when fully populated. Against a 170Mi limit, that is a meaningful fraction.

  6. Quantify the snapshot. Count Services and Endpoints. The official scaling guidance estimates required memory as (Pods + Services) / 1000 + 54 MB for a default deployment, and (Pods + Services) / 250 + 56 MB with autopath. If your measured steady state is well above the estimate, something else (cache, goroutines, leak) is contributing. If it matches and your limit is below the estimate plus re-list headroom, the limit is simply wrong.

Metrics and signals to monitor

SignalWhy it mattersWarning sign
process_resident_memory_bytes vs. container limitThis is what the OOM killer usesRSS > 70% of limit: investigate. > 85%: critical
Post-GC heap minimum trendLeak detection independent of Go’s spiky allocationMinimum rising steadily, never reclaiming
go_goroutinesBlocked calls and leaks accumulate here firstSustained > 2x baseline; growth without QPS growth
go_gc_duration_secondsGC pauses lengthen as heap pressure growsPauses > 10ms; increasing trend
coredns_cache_entries vs. configured maxCache at capacity is both a symptom and a memory consumerPinned at max with rising evictions
coredns_proxy_request_duration_seconds{to=...}Slow upstreams cause the goroutine accumulation that becomes memory pressureP99 > 250ms for any upstream
Pod restart count and OOMKilled eventsThe cliff edge itselfAny restart with OOMKilled reason

Fixes

Limit too small for the re-list peak

Increase the memory limit so the startup re-list peak fits with headroom. The re-list can reach 2-3x steady state in large clusters, so a limit 10-20% above steady state is a crash loop waiting for its first restart. This is the correct emergency fix during an active loop: kubectl edit deployment coredns -n kube-system and raise the limit. It restarts the pods, but they are already crash looping, so you are not making anything worse. Size the new limit from the scaling formula plus re-list headroom, not from whatever number was there before.

Cluster grew past the limit

Same fix, different root cause. Re-run the sizing estimate with current Pod and Service counts and set the limit accordingly. Then alert on the growth rate: linear extrapolation of steady-state RSS against the limit gives you runway in days, and you want to re-size before you arrive, not after.

Cache too large for the limit

Reduce the cache size in the Corefile, but understand the tradeoff: a smaller cache means more evictions, a lower hit ratio, more upstream load, and higher latency. The better resolution is usually to raise the limit to fit the cache your working set needs, using cache entries and eviction rate as the sizing evidence. Shrinking the cache to fit an undersized limit trades one failure mode (OOM) for another (upstream amplification).

Goroutine or connection leak

Confirm with the go_goroutines trend. If goroutines accumulate when upstreams are slow, the fix is at the upstream or network layer, and max_concurrent on the forward plugin gives you backpressure (rejected queries get REFUSED instead of unbounded goroutine growth; each in-flight query costs roughly 2KB of memory). If the count grows independent of traffic, suspect a plugin bug: connection and goroutine leaks have been fixed across releases, so check your CoreDNS version against release notes and upgrade. For a persistent unidentified leak, enable pprof and capture a heap and goroutine profile to find the allocation source.

Go runtime behavior

If your CoreDNS is built with Go 1.19 or later, setting the GOMEMLIMIT environment variable to just below the container limit makes the GC work harder as RSS approaches the cap, at the cost of CPU. This is a mitigation, not a fix: in a GC death spiral you trade an OOM kill for CPU saturation. It buys time and smooths the approach to the cliff, but correct limit sizing is still the answer.

Prevention

  • Size for the peak, not the mean. Set the limit to cover the startup re-list peak: 2-3x steady-state RSS in large clusters, with a floor from the scaling formula. Keep steady-state RSS under 70% of the limit.
  • Alert on the approach, not the corpse. OOMKilled is a post-mortem signal. Alert on RSS > 70% of limit (investigate) and > 85% (critical), on post-GC heap minimum trending up, and on goroutine growth without QPS growth. These are the signals that exist while you can still act.
  • Track cluster growth against the limit. Service and Endpoint counts are leading indicators of snapshot size. When the cluster grows, re-size before the first OOMKill, not after.
  • Re-check sizing after topology changes. Rollouts that restart all CoreDNS pods simultaneously re-trigger the re-list peak on every pod at once and also cold-start every cache. Stagger rollouts (maxUnavailable=1, PodDisruptionBudget) so you never test the re-list peak cluster-wide.
  • Keep CoreDNS current. Memory-relevant fixes (connection leaks, socket caps, watch handling) ship in patch releases. Staying current removes known leak sources from the suspect list.
  • Mind the apiserver restart coupling. The re-list peak also fires when the apiserver restarts and watches reconnect. After control plane maintenance, watch CoreDNS RSS the same way you watch the apiserver.

How Netdata helps

  • RSS against the limit, per pod. Netdata charts process_resident_memory_bytes alongside container memory limits from cgroup data, so the approach to the cliff is visible as a percentage per CoreDNS replica, not as an aggregate that hides one dying pod.
  • Goroutine and heap correlation. Plotting go_goroutines, go_memstats_heap_inuse_bytes, and GC pause duration on the same timeline separates a leak (all climbing together) from a one-off spike (heap reclaims, goroutines return to baseline).
  • Cache pressure signals. coredns_cache_entries and coredns_cache_evictions_total per pod show whether the cache is the memory consumer and whether it is undersized for the working set.
  • Restart and OOMKill context. Pod restart counts and container state, correlated with the memory trend leading into each kill, tell you immediately whether you are looking at a slow leak or the re-list crash loop.
  • Upstream latency on the same dashboard. Per-upstream forward latency next to goroutine count confirms or rules out slow-upstream accumulation as the memory driver, without switching tools.